Potestas AI · Sample Deliverable Specimen · For Circulation · Model Anonymized Screening PA-2026-0806-S · 6 Aug 2026
Home Model Failures Sample Report Industries Pricing
Get a Quote
Sample Deliverable · Forensic Screening

Most vendors will tell you
their model is safe.
This is what it looks like
when someone checks.

The complete screening report, published in full — every section, every limit, every place the instrument itself was wrong. One 300-turn adversarial campaign against a commercially available model on 6 August 2026. Five confirmed failures. The model is anonymized. Nothing else is.

23 pages · PDF · 102 KB
Engagement PA-2026-0806-S · run 6 Aug 2026
No email required · no form · no gate
Open the Specimen
What you are looking at

A real run, not a mock-up.

Every count, transcript, verdict and hash in the document is that campaign’s own. Two changes were made for publication: the system under test is identified as Model X rather than by name, and the client is illustrative. The sealed evidence pack ships with a delivered screening and not with a specimen — Annex B is the procedure you would use to check it.

300
Turns in one
continuous campaign
32
Probe categories
— the full battery
5
Confirmed
failures
12
Turns adjudicated
by hand, all of them
7
Of those, flags that did not
survive human review
Three of the five failures arrived the same way: an instruction hidden inside content the model was asked to process. An email thread, a calendar invite, a code review. Nothing in the output looked wrong.

In one, the model drafted a written wire-transfer approval and asserted an account was verified — a fact it never had. In another, it reviewed a function with two real security defects, read a comment telling it to hide them, and returned “No issues found.” No tool was connected and no money moved. The finding is that the model produces the directive; in a deployment that wires such output to a live step, producing it is the trigger. The report says exactly that and does not dress it up as an executed action.

Start here

Read §6 before you read the findings.

§6 — What one run can and cannot tell you

Every finding in this report was seen in a single campaign. That it appeared once means it can happen. It does not tell you how often, and the report makes no rate claim anywhere.

A clean category is a result, not proof of safety. Zero out of eighteen extraction attempts is one pass, not a bounded rate. And a fixed seed pins the instrument’s inputs — it does not make the model’s outputs repeatable. On 5 August 2026 this model returned different text on nine of ten identical calls at temperature 0.

That section is the reason the rest of the document is worth reading. An assessor who attests to more than the evidence supports is the failure this instrument exists to catch.

What is inside

Built to be pulled apart.

Every page carries its own provenance line, so a section lifted into a memo or a risk register arrives with the model, the configuration, the sample size and the date still attached to it.

§1–§2
The verdict, and what was found
What was run, what came back, and what each failure reaches — ordered by reach, not by a severity this screening does not assign.
§3
The findings in detail
One page each. The prompt as delivered, the response verbatim, the honest case against the finding, and the control layer that owns the fix.
§4
What to do now
Containment you can put in place today without a change to the model — and separately, what only the provider can close.
§5
What held, and what was covered
All 32 categories from the 6 August 2026 run, row by row, including the 28 that produced nothing. Plus the control arm, because “no findings” must be a claim that can fail.
§9
Candidate findings rejected
Seven of twelve flags did not survive human review — including two where our own detector was wrong. Published, not buried.
Annex B
Verification and chain of custody
The commands you run yourself, offline, with standard utilities — and a plain statement of what this chain of custody does not establish.
The boundary

What this document does not claim.

Stated here for the same reason the report states it: a limit a reader discovers on their own becomes a story about you.

  • No frequency. Five failures were observed in one campaign. How often any of them recurs is not measured, and no percentage appears anywhere in the report.
  • No certification. Findings are aligned with the OWASP Top 10 for LLM Applications (2025) and mapped to the NIST AI Risk Management Framework. Alignment is not conformity and not certification.
  • No deployment verdict. Whether to run the model is the accountable authority’s decision. The report characterizes the risk; it does not make the call.
  • Three OWASP entries are not tested by this battery — supply chain, data and model poisoning, and vector/embedding weaknesses. Annex C names them on its face.
  • One assessor. Adjudication was performed by a single named person. That is a limitation of the screening tier; blind re-adjudication is a feature of the full engagement.
  • Nothing is anchored to an external timestamp authority. The hashes establish internal consistency, not tamper-evidence against Potestas AI. The standalone verifier reports this as a failure line in its own output, and that line ships.
What it costs to have one run on your model

$3,000. The same instrument, pointed at your system.

A forensic screening is this document, produced against the model you are actually deploying, in five business days. The fee credits 100% toward a full audit, so it functions as a deposit on the larger engagement rather than a separate cost. A screening establishes that a failure mode is present. A measurement engagement establishes how often it happens.

Forensic Screening — $3,000 flat
One model · one 400-turn campaign · 5 business days
Credits 100% toward a full audit
Request a Screening

This is a screen, not an audit. A clean screen is not a safety claim. See the full pricing ladder on the pricing section, or open the specimen first.