Potestas AI Forensic Screening · $3,000 · Self-Serve 833-TEST-LLM
Home Model Failures Sample Report Industries Pricing
Buy the Screen
Entry · Screening

What it did.
Not what it said.

The agentic say-vs-do screen. $3,000 flat, delivered same day.

Your agent can move money, send mail, reach records, and call other systems. When it does something it should not have, the first thing you get is its own account of what happened — and a model will calmly describe refusing a request it actually carried out.

This screen hands that model a hundred situations and live tools, then reads the tool-call log: what it really invoked, and with what arguments. The distance between those two things is where the damage lives.

You give us the name of a model. No API key, no credentials, no system prompt, no access to anything you run. We use our own access. Your entire exposure is the $3,000.
Buy the Screen — $3,000 → See a Finished Report — 22 pages
What leaves your organization
The complete list. Nothing is abbreviated.
API keyNONE
Credentials or tokensNONE
System promptNONE
Tool definitionsNONE
Network or system accessNONE
Customer or business dataNONE
The model’s nameTHAT IS IT
Same day, in returnREPORT · EVIDENCE PACK
Why this exists

Everything you know about your agent
comes from the agent.

Your logs record what it reported. Your incident review opens with its account of itself. That account is the least reliable artifact in the whole system, and it is the one everything downstream is built on.

What it says
A fluent, reasonable sentence.
Written to be read by a person, often describing exactly the judgement you would have wanted it to make. This is the part that gets pasted into the postmortem.
What it does
A function call, with arguments.
No account of itself, no framing, no intent. Just the call it made and what it passed. The money moves on this line, not on the sentence.
We ran exactly this against four frontier agents and published the whole thing — method, transcripts, limits. Two of them wired the money.
Read the Study →
What you are risking

A purchase that needs a security review
is not a $3,000 purchase.

Most vendors need access to your model before they can test it. That one requirement is what turns a quick decision into a three-month one. Below is every step it normally triggers, and why none of them reaches you here.

Security reviewNothing to review
API key provisioningWe use our own
Vendor access onboardingNo access granted
Data processing agreementNo data processed
InfoSec sign-offNo system touched
Legal review of data handlingNo data handled
Competitive solicitationUnder the micro-purchase threshold
Sending us your system promptThere is no field for one
What remains is a card, a model name, and an afternoon. Nothing to install, nobody to onboard, no access to grant, no meeting to schedule — and your answer back the same day.
If the screen finds nothing
You spent $3,000 establishing that the model held across a hundred adversarial situations with live tools in its hands. We will tell you that plainly, and tell you what the full battery tests that a screen does not.
If the screen finds something
You spent $3,000 finding it before your customers, your regulator, or a journalist did — with the tool-call log that proves it, in a format built to survive someone else's scrutiny.
What it establishes

Presence. Never frequency.

Two different questions with two different prices. Here is the one $3,000 answers — and the one it does not.

This screen answers
Did it happen at all?
One pass over 100 scenarios: it took these actions when handed a tool. Presence is a count — the scenarios that breached, and the calls that constitute the breach. That is a finding you can put in front of someone. It is not a measurement.
The full stress test answers
How often?
A behaviour seen once may be a 2% rate or a 60% rate, and those demand different responses. A presence finding cannot be sized, priced, or remediated against. Producing a rate takes independent re-runs and a confidence interval — a separate engagement.
No rate, no confidence interval, and no cross-model comparison is produced at this tier. If a screen surfaces something you need sized, this fee credits 100% toward the Forensic Stress Test — you do not pay twice to ask the second question.
How it is graded

Off the log. No judge model.

No second model reads the transcript and decides whether something bad happened.

The scenario hands the model working tools. The grader reads which ones it called, and with what arguments. A model that narrates a wire transfer it never executed did not execute it, and the log says so. That gap — between what a model says and what it does — is the entire object of this screen.

It is also why you get it back today. No judge model means no review queue and about a minute of compute. The speed comes from how it is graded, not from anyone rushing your job through.

The text battery is judged, human-adjudicated on every queued turn, and takes days. That is why it is not in this tier. The split follows the instrument, not the price list.
What you receive

Not a score. A record.

Every finding in the report has the same anatomy, and every part of it is evidence rather than assertion.

Anatomy of one findingFormat, not an example
The scenario
Which of the hundred, by identifier, and the category of failure it probes. Fixed seed, so it can be run again.
What it was handed
The situation put to the model and the tools made available to it — the exact conditions, reproducible.
What it said
The model's own words, in full. Often reasonable. Sometimes a refusal.
What it called
The tool invocation, with its arguments, straight from the log. This is the finding. Everything else is context around it.
The gap
The sentence beside the call it contradicts, on the same page. This is the finding you can hand to someone who was not in the room — not a claim that the model misbehaved, but the two artifacts side by side with the distance between them visible.
Chain of custody
Sealed pack, per-run manifest, replay record. Built so that someone who does not trust us can check it.
The deliverable

Do not take our word for the quality.
We published an entire engagement.

Not a sanitised sample, not a redacted excerpt, not a template with the numbers pulled out. A finished 22-page report, on the public internet, with its own limits printed inside it. Read it before you spend anything — if it does not convince you, this page should not either.

Published report · 22 pagesRead it free, right now
§1–§2
The verdict, and what was found
Ordered by what each failure actually reaches, not by a severity number.
§3
The findings in detail
One page each. The prompt as delivered, the response verbatim, and the honest case against the finding.
§4
What to do now
Containment you can put in place today without touching the model — and, separately, what only the vendor can fix.
§5
What held, and what was covered
Every category, row by row — including the ones that produced nothing. A report that only lists failures is a sales document.
§9
Candidate findings we rejected
Seven of twelve flags did not survive human review — including two where our own detector was wrong. Printed in the deliverable, at our own expense. Ask any other vendor for this section.
Annex B
Verification and chain of custody
The commands you run yourself, offline, with standard utilities — plus a plain statement of what this chain of custody does not establish.
One thing to be precise about, since the whole page is built on precision. The published report is a text-battery engagement — no live tools were connected to that run. Your $3,000 screen is the agentic axis: same format, same standard, same refusal to hide the limits, but the findings are tool calls rather than turns. We are showing you the workmanship you are buying, not a copy of the document you will receive.
The honest limit

Where a human still reads it.

Stated here, rather than discovered afterwards.

The agentic axis usually seals without adjudication. A turn can still queue — when a model narrates an action its log does not show, or when a detector flags a refusal that was actually correct. We have measured both.

When it happens the scenario goes to a human, not to a guess. It is resolved by a standing hand-audit, it adds no rate, and it is included in the fee — returned within two business days of delivery.

What you are buying is presence — not a promise that nothing ever needs a human. Anyone selling you that promise is overselling.

Before you buy

The five questions worth asking.

Answered here so you do not have to book a call to find out.

How is this different from a benchmark score?
A benchmark grades the sentence. This grades the call. A model that writes a perfectly worded refusal and then invokes the tool anyway scores well on a benchmark and fails here — and once your agent has tools in production, that is the only one of the two that can cost you anything.
You are testing the public model, not my deployment. Isn't that a different thing?
Yes, and the report says so in as many words rather than letting you assume otherwise. What $3,000 buys is a clear answer about what you are building on. What your own system prompt and tool set do on top of it is a real and separate question — arranged directly, never through a web form, and priced accordingly.
What if you find nothing?
We tell you plainly, and tell you what the full battery covers that a screen does not. There is no guarantee at this tier — no-findings-no-fee applies to the full battery only. And a clean screen is not a safety claim: it means these 100 scenarios did not surface a breach in one pass. Worth knowing, and not the same as safe.
Why should I trust your result?
You shouldn't have to. Every finding arrives with the tool-call log that produced it, inside a sealed pack with a manifest and a replay record. The grading is deterministic, so there is no model judgement to take on faith. And we publish our own report in full, including where the instrument was wrong.
Who is actually accountable for this?
A named operator with an active clearance, who signs the work. Not a dashboard, not an anonymous crowd, and not a queue. If a scenario escalates, the person who hand-audits it is the person whose name is on the report.
What this is not

The limits, in full.

Every one of these is something a screen genuinely cannot do. Listed so that nobody discovers one after paying.

Purchase

Four steps, start to delivered.

Nothing to install, nobody to onboard, no access to grant.

01

Pay

$3,000 by card, through Stripe. Receipt immediately. Card details go to Stripe and never touch this site.

02

Name the model

On the form you land on straight after payment. A name, an email, and the model you want screened. Under a minute — and nothing sensitive leaves your side.

03

We run it

One pass of the 100-scenario suite through our own access to the model. Current instrument version, fixed seed, tool-call log captured per scenario. About a minute of compute.

04

Delivered the same day

Findings report and sealed evidence pack by email. If a scenario escalated to a hand-audit, that result follows within two business days — included.

You already know what your agent says it did.
This is the only way to find out what it did.
If you would rather talk first, notice what there would be to talk about. No access to negotiate, no data to scope, no integration to plan — the only open question is whether the answer is worth $3,000 to you. Email the founder directly and you will get a person, not a form response.
$3,000
Flat · one model · one pass
Access we needNone
You supplyA model name
Scenarios100 agentic
GradingTool-call log
DeliverySame day
Escalated turnsIncluded
Credits toward audit100%
Card purchase not yet live
Card purchase is not live on this page yet.
Email to purchase and an invoice comes back the same day.
Already paid? Name your model →

You are going to be asked this question eventually.

How do you know the model you built on does what it says? It will come from a board, a regulator, a customer's security team, or a court — and the honest answer, today, is that you are taking the model's word for it.

$3,000 stops you having to. One pass, a hundred situations, and a log of what it actually did — back today.

Buy the Screen — $3,000 → Compare the Tiers