The complete screening report, published in full — every section, every limit, every place the instrument itself was wrong. One 300-turn adversarial campaign against a commercially available model on 6 August 2026. Five confirmed failures. The model is anonymized. Nothing else is.
Every count, transcript, verdict and hash in the document is that campaign’s own. Two changes were made for publication: the system under test is identified as Model X rather than by name, and the client is illustrative. The sealed evidence pack ships with a delivered screening and not with a specimen — Annex B is the procedure you would use to check it.
In one, the model drafted a written wire-transfer approval and asserted an account was verified — a fact it never had. In another, it reviewed a function with two real security defects, read a comment telling it to hide them, and returned “No issues found.” No tool was connected and no money moved. The finding is that the model produces the directive; in a deployment that wires such output to a live step, producing it is the trigger. The report says exactly that and does not dress it up as an executed action.
Every finding in this report was seen in a single campaign. That it appeared once means it can happen. It does not tell you how often, and the report makes no rate claim anywhere.
A clean category is a result, not proof of safety. Zero out of eighteen extraction attempts is one pass, not a bounded rate. And a fixed seed pins the instrument’s inputs — it does not make the model’s outputs repeatable. On 5 August 2026 this model returned different text on nine of ten identical calls at temperature 0.
That section is the reason the rest of the document is worth reading. An assessor who attests to more than the evidence supports is the failure this instrument exists to catch.
Every page carries its own provenance line, so a section lifted into a memo or a risk register arrives with the model, the configuration, the sample size and the date still attached to it.
Stated here for the same reason the report states it: a limit a reader discovers on their own becomes a story about you.
A forensic screening is this document, produced against the model you are actually deploying, in five business days. The fee credits 100% toward a full audit, so it functions as a deposit on the larger engagement rather than a separate cost. A screening establishes that a failure mode is present. A measurement engagement establishes how often it happens.
This is a screen, not an audit. A clean screen is not a safety claim. See the full pricing ladder on the pricing section, or open the specimen first.