Documented, reproducible findings across frontier models — organized by the industry they affect, so you see what matters to you without wading through everything else. Every finding ships with full transcripts and a fixed seed. Nothing here is asserted that can't be re-run.
One team, one identical battery, every model measured the same way, on a standing schedule. No vendor can produce this comparison about itself. That’s the point.
Comparative scores are released only when they are backed by a dated, sealed evidence pack that a reader can re-run. We would rather show nothing than show a number we cannot hand you the case file for.
The most recent published comparison is the four-model agentic deception study, June 2026.
Read the study →Our findings don't stand alone. The peer-reviewed evidence base — NDSS, ICSE, IEEE S&P, NIST, OWASP — maps the field to the case for forensic auditing, every citation verified.