Your agent can move money, send mail, reach records, and call other systems. When it does something it should not have, the first thing you get is its own account of what happened — and a model will calmly describe refusing a request it actually carried out.
This screen hands that model a hundred situations and live tools, then reads the tool-call log: what it really invoked, and with what arguments. The distance between those two things is where the damage lives.
Your logs record what it reported. Your incident review opens with its account of itself. That account is the least reliable artifact in the whole system, and it is the one everything downstream is built on.
Most vendors need access to your model before they can test it. That one requirement is what turns a quick decision into a three-month one. Below is every step it normally triggers, and why none of them reaches you here.
Two different questions with two different prices. Here is the one $3,000 answers — and the one it does not.
No second model reads the transcript and decides whether something bad happened.
The scenario hands the model working tools. The grader reads which ones it called, and with what arguments. A model that narrates a wire transfer it never executed did not execute it, and the log says so. That gap — between what a model says and what it does — is the entire object of this screen.
It is also why you get it back today. No judge model means no review queue and about a minute of compute. The speed comes from how it is graded, not from anyone rushing your job through.
Every finding in the report has the same anatomy, and every part of it is evidence rather than assertion.
Not a sanitised sample, not a redacted excerpt, not a template with the numbers pulled out. A finished 22-page report, on the public internet, with its own limits printed inside it. Read it before you spend anything — if it does not convince you, this page should not either.
Stated here, rather than discovered afterwards.
The agentic axis usually seals without adjudication. A turn can still queue — when a model narrates an action its log does not show, or when a detector flags a refusal that was actually correct. We have measured both.
When it happens the scenario goes to a human, not to a guess. It is resolved by a standing hand-audit, it adds no rate, and it is included in the fee — returned within two business days of delivery.
What you are buying is presence — not a promise that nothing ever needs a human. Anyone selling you that promise is overselling.
Answered here so you do not have to book a call to find out.
Every one of these is something a screen genuinely cannot do. Listed so that nobody discovers one after paying.
Nothing to install, nobody to onboard, no access to grant.
$3,000 by card, through Stripe. Receipt immediately. Card details go to Stripe and never touch this site.
On the form you land on straight after payment. A name, an email, and the model you want screened. Under a minute — and nothing sensitive leaves your side.
One pass of the 100-scenario suite through our own access to the model. Current instrument version, fixed seed, tool-call log captured per scenario. About a minute of compute.
Findings report and sealed evidence pack by email. If a scenario escalated to a hand-audit, that result follows within two business days — included.
How do you know the model you built on does what it says? It will come from a board, a regulator, a customer's security team, or a court — and the honest answer, today, is that you are taking the model's word for it.
$3,000 stops you having to. One pass, a hundred situations, and a log of what it actually did — back today.