Potestas AI Model Failures · Documented · Reproducible · Updated on a Standing Schedule 833-TEST-LLM
Home Model Failures Sample Report Industries Pricing
Get a Quote
Model Failures · The Standing Record

Where frontier AI
actually breaks.

Documented, reproducible findings across frontier models — organized by the industry they affect, so you see what matters to you without wading through everything else. Every finding ships with full transcripts and a fixed seed. Nothing here is asserted that can't be re-run.

Current Intel

The state of play,
on the record.

One team, one identical battery, every model measured the same way, on a standing schedule. No vendor can produce this comparison about itself. That’s the point.

Get the full briefing — every model, every industry — in your inbox.
Joined by CISOs, procurement officers, and safety teams. No spam. Unsubscribe anytime.
Internally, we call this the AI COP — the Common Operating Picture on how frontier models are actually behaving. Old habits.
Integrity RankingsNext Cycle
Rankings publish on completion of the current battery.

Comparative scores are released only when they are backed by a dated, sealed evidence pack that a reader can re-run. We would rather show nothing than show a number we cannot hand you the case file for.

The most recent published comparison is the four-model agentic deception study, June 2026.

Read the study →
Pick Your Industry
Select an industry — each finding re-summarizes for your world.
KATANA-RR-2026-01 · Flagship Study Say-vs-Do · Agentic Testing

We tried to socially engineer four frontier AI agents. Two of them wired the money.

4-Model Comparative Study · 9 Agentic Scenarios · Published June 2026

KATANA-2025-002 · Critical Finding Reasoning-Integrity Validation

The AI faked its own reasoning — and still got the right answer.

Chain-of-Thought Fabrication · Grok-4 observed · other models under test

KATANA-2025-003 · Critical Finding Self-Report vs. Actual Behavior

The AI's own safety check said "all clear" — while it was lying.

The Liar's Protocol · Performative Compliance · Grok-4 primary · other models confirmed

KATANA-2025-001 · High Severity Sustained-Use · Conversational Testing

Push the AI long enough, and its basic math quietly breaks.

Prime Number Bias Under Sustained Pressure · Grok-4 observed

No published findings for this industry yet — new findings land on a standing ~10-day schedule. Check back, or get the briefing above.
This record grows on a standing schedule. New stress tests run continuously against every frontier model — including every new release — and findings publish roughly every ten days. One-shot tools give you a snapshot. This is the only continuously-updated integrity record of how these models change over time.

The research behind the method.

Our findings don't stand alone. The peer-reviewed evidence base — NDSS, ICSE, IEEE S&P, NIST, OWASP — maps the field to the case for forensic auditing, every citation verified.

Open the Evidence Base →