Potestas AI Screening Intake · Step 2 of 4 833-TEST-LLM

Name the model.

That is the only thing left to do.

No API key, no credentials, no system prompt, no access to anything you run — there is nowhere on this form to put them, on purpose. Tell us which model you are building on and we run it through our own access. Your screen runs today.

What lands in your inbox tonight

Three artifacts. All of them evidence.

Not a dashboard, not a score, and not a percentage. A record built so that somebody who does not trust you — or us — can check it themselves.

01 · The report
Every scenario that broke, and exactly how
What the model was handed. What it said. What it actually invoked, with the arguments. And the gap between the last two — the sentence that promised one thing and the call that did another. Written to be read by a board, not by an engineer.
Findings report · PDF
02 · The record
All one hundred scenarios, not just the failures
Every tool invocation across the whole run, with its arguments, straight from the log. The clean scenarios are evidence too — they are what let you say the model held, and point at the reason rather than at a percentage.
Tool-call log
03 · The seal
A pack that survives someone else’s scrutiny
All artifacts in one archive under a SHA-256 manifest, with the replay record and the exact instrument version stamped in. This is the part that matters in a procurement review, an audit, or a room where somebody is asking how you know.
Sealed evidence pack · ZIP

And you can check it without us.

The pack ships with the verification steps written out — commands you run yourself, offline, with standard utilities, plus a plain statement of what the chain of custody does not establish. Fixed seed and a named instrument version mean the exact conditions can be re-run: by you, by your auditor, by a vendor arguing with the result, or by us in six months when you want to know whether anything changed.

That is the difference between evidence and an assurance. An assurance asks you to trust the person giving it. This does not.

Seed fixed
Instrument version stamped
Artifacts hashed
Conditions re-runnable
Verification offline
And the fee is not spent
$3,000what you paid today
−$3,000credited in full toward a Forensic Stress Test
$0net cost, if this becomes a full engagement
Which leaves two outcomes, and both of them are good ones. The screen finds something, you take it further, and today's fee comes off that engagement entirely. Or it finds nothing, and $3,000 bought you a clean, evidenced answer about the model your product is built on — which is roughly what one afternoon of the meeting you are avoiding would have cost.
Under a minute

Tell us the model.

Online submission is not live on this page yet.
Email the model name to joseph.cirello@potestasai.com and your screen runs the same day.
Section 1 · Where the report goes
So the deliverable reaches the right person, and so we can match this to your payment.
Optional, but it speeds things up. Without it we match on your email address instead.
Section 2 · The model
One model, one pass. This is the only thing we actually need from you.
The precise string, including the version or date suffix if it has one. A finding is bound to a specific build, so an approximate name makes the record less useful to you.
For a model from OpenAI, Anthropic, Google or xAI, leave this blank — we already have access and run the suite ourselves. For anything else — open-weight, self-hosted, a smaller or regional vendor, or served through a third-party host — say where it lives and we confirm we can reach it before we take the run.
Section 3 · What we screen
Pick one. The first needs nothing further from you, and it is what most people choose.
⚠  Do not paste a system prompt anywhere on this form. There is deliberately no field for one. If the second option applies to you, we agree a channel first — you should not be sending proprietary configuration to a vendor you have not spoken to.
Section 4 · Anything else
Optional. Skip it and we run the standard suite.
Useful, but keep it non-sensitive — this box goes through the same ordinary web plumbing as the rest of the form.
✓  Received. Your screen runs today and the evidence pack comes back to the email above.
Questions in the meantime: joseph.cirello@potestasai.com

From here

Today
The 100-scenario suite runs against the model. Fixed seed, current instrument version. About a minute of compute.
Today
Findings report and sealed evidence pack, by email, to the address you give above.
2 days
If a scenario escalated to a hand-audit, that result follows — included in the fee, never guessed.

Said plainly

Presence, not frequency. This establishes whether a behaviour occurred. It does not establish how often, and no rate appears in the deliverable.

About the model, not your deployment. Screening what the vendor ships tells you what you are building on. It cannot tell you what your own prompt and tools do on top — and the report says so rather than letting you assume otherwise.

A clean screen is not a safety claim. If nothing surfaces we will say so directly, and explain what the full battery tests that a screen does not.