Potestas AI · Federal & Defense SAM.gov Registered · NAICS 541715 · Secret Clearance Active Disabled Veteran-Owned · Providence RI
Home Model Failures Sample Report Industries Pricing
Get a Quote
Regulated & High-Assurance · Every Industry

When the wrong answer
has consequences.

Potestas AI is a disabled veteran-owned small business providing forensic LLM stress testing for federal, defense, and high-assurance deployments. Every engagement produces a cryptographically sealed evidence package legally defensible in any proceeding.

ItemStatus / ValueNotes
SAM.govRegistered ✓EIN: 41-2476590 · Active registration
NAICS541715Research & Development in Physical, Engineering, and Life Sciences
Veteran StatusDisabled Veteran-OwnedService-disabled veteran owned and operated
Secret ClearanceActiveJoseph Cirello · Facilitates classified deployment contexts
Business TypeDisabled Veteran-Owned Small BusinessProvidence, Rhode Island
Air-Gap CapabilityAvailableAdversarial engine runs fully on-premises · Zero data egress
ITAR HandlingAvailable by arrangementITAR-sensitive deployment auditing available by arrangement
A model that hallucinates a bullet count of 79 when the truth is 30 cannot be trusted in mission-critical logistics. Katana is the safety interlock — not a replacement for the LLM.
— Potestas AI · DoD Application Statement
Why This Matters for Federal Deployment
Chain-of-Custody
Legally Defensible Evidence
Every engagement produces a FINGERPRINT.DB — per-turn cryptographic hash plus semantic signature. SHA-256 manifest seals all artifacts. Holds up to IG review, legal proceedings, and procurement audits.
Reproducibility
TEVV-Aligned Fixed-Seed Protocol
Fixed random seed (42) at temperature 0.0. Every audit run is fully reproducible for independent verification — regulatory-grade auditability aligned with DoD TEVV standards.
Zero Trust
LLM as Untrusted Component
Katana treats the LLM as an untrusted component requiring behavioral governance at every output. Verified forensic evidence of integrity — not blind trust in vendor assurances.
Air-Gap
Fully On-Premises
Our adversarial engine runs locally, fully air-gapped. Zero network connectivity required once deployed. No data leaves the network — a requirement most regulated and defense environments demand.
Independent Judging
The Auditee Never Grades Itself
The model being audited is never used to evaluate its own responses. Cross-family LLM judge eliminates evaluator bias and produces independent forensic verdicts.
Severity Weighting
Weighted Robustness Index
Weighted Robustness Index — severity-weighted audit score (0–100%) reflecting actual risk exposure. Probe severities range from 1 (noise) to 10 (critical jailbreak).
Compliance Framework Alignment

Mapped to the Frameworks Your Auditors Already Use.

Every finding is reported in the language your compliance, risk, and procurement teams already work in — so the evidence drops straight into the documentation you're required to produce.

Risk Management
NIST AI RMF 1.0
Evidence package — severity scoring, patch recommendations, chain-of-custody — directly supports GOVERN, MAP, MEASURE, and MANAGE functions of the NIST AI RMF.
Aligned ✓
Vulnerability Classification
OWASP LLM Top 10 (2025)
Probe categories map directly to OWASP LLM Top 10 including LLM01 (Prompt Injection), LLM07 (System Prompt Leakage), and LLM10 (Unbounded Consumption).
Aligned ✓
AI Management Systems
ISO/IEC 42001
Reproducible audit methodology and structured evidence package support ISO/IEC 42001 AI management system requirements — risk assessment, monitoring, and continual improvement documentation.
Aligned ✓
DoD Evaluation Standard
TEVV — Test, Evaluation, Verification & Validation
Fixed seed (42), temperature 0.0, deterministic run configuration. Katana audits are fully reproducible for independent verification — the core requirement of DoD test and evaluation. TEVV-aligned by design, with a sealed evidence trail any independent evaluator can replay.
Aligned ✓
Adversarial Threat Taxonomy
MITRE ATLAS
Katana probe categories map to MITRE ATLAS adversarial tactics and techniques for AI systems — the threat-informed framework increasingly referenced in federal AI security requirements. Findings are reported in language that maps to the ATLAS matrix.
Aligned ✓
Agentic Risk
OWASP Agentic AI Top 10
Our flagship agentic study targets Agent Goal Hijack and Tool Misuse — the two leading categories in the OWASP Top 10 for Agentic Applications. As AI agents gain authority over funds, credentials, and systems, this is the threat surface that matters most.
Aligned ✓
EU Regulation
EU AI Act
Forensic evidence package supports high-risk AI system conformity assessment documentation requirements for general-purpose AI model evaluation.
Aligned ✓
Adversarial Testing
NIST AI 100-2 E2025
Probe library reflects current adversarial ML research — CRESCENDO_ESCALATION (USENIX 2025), MANY_SHOT_INJECTION (Anthropic MSJ), INDIRECT_INJECTION (OWASP LLM01).
Aligned ✓
The Regulatory Floor Is Rising

Adversarial Testing Is Moving From Optional to Expected.

The shift is underway across regulators, insurers, and standards bodies. The organizations that build a defensible behavioral-evidence record now are the ones that won't be scrambling when it becomes a baseline requirement. The dates below are current as of June 2026 — and they are moving in one direction.

EU AI Act
Adversarial testing required for high-risk systems
High-risk obligations were deferred to December 2027 (Annex III) under the May 2026 Digital Omnibus agreement. Transparency obligations remain active from August 2026. The mandate is delayed — not removed.
NIST · Federal
AI RMF + GenAI Profile are the federal baseline
The AI RMF MEASURE function calls for systematic evaluation of AI behavior. A critical-infrastructure profile concept note followed in April 2026, extending the framework toward regulated sectors.
Insurance · NAIC
Behavioral underwriting is arriving
ISO AI liability exclusions took effect January 2026. Over half of U.S. states have adopted the NAIC AI model bulletin; a third-party AI oversight model law is anticipated in 2026.
The market is already moving without waiting for the regulators.
Independent industry analyses place the AI red-teaming services market in the low-single-digit billions in 2025, projected to grow many-fold over the next decade. Multiple 2026 surveys report that a majority of enterprises still rely on traditional, non-adversarial testing — and that breaches involving compromised AI agents carry measurably higher remediation costs than conventional incidents. The gap between deployment and verification is the exposure. Forensic auditing closes it with evidence, before it becomes a finding.
Regulatory dates reflect publicly reported developments as of June 2026 and are subject to change through formal adoption. Market figures are drawn from third-party industry analyses and are directional, not guarantees. Potestas AI states alignment with the frameworks above — not third-party certification.
Procurement Language

What to Put in Your SOW.

The terms below describe Katana's capabilities accurately for Statements of Work, contract requirements, and procurement language.

TermDefinition for Procurement Use
Forensic AI AuditRepeated sustained adversarial campaigns plus a measurement phase, producing a measured failure rate with a confidence interval for every confirmed finding, inside a cryptographically sealed evidence package with chain-of-custody documentation.
Deep-Hop ProtocolMulti-turn adversarial pressure campaign where each turn compounds on prior failures, reaching 8–13 hop depth in Deep Compute phase.
Air-Gapped Adversarial AgentAdversarial engine operating fully on-premises. Zero network egress after deployment. Air-gappable by design.
Weighted Robustness Index (WRI)Severity-weighted aggregate integrity score (0–100%) reflecting actual risk exposure across 36 probe categories.
FINGERPRINT.DBPer-turn cryptographic response hash plus semantic signature. Tamper-evident chain-of-custody artifact. Verifiable by any third-party auditor.
Fixed-Seed ReproducibilityRandom seed fixed at 42, temperature 0.0. Every run fully reproducible for independent verification — TEVV-aligned regulatory-grade auditability.
Small Business StatusPotestas AI is a Disabled Veteran-Owned Small Business, registered and active in SAM.gov under NAICS 541715.
Published Research · Katana Corpus

Named Findings That Changed
How We Evaluate AI.

The Katana audit corpus has produced named, documented, reproducible vulnerability findings across every major frontier model. These are not theoretical — they were confirmed in field evaluations including an authorized US Army deployment.

KATANA-RR-2026-01 · Flagship Study
Agentic Deception: Two of Four Models Wired the Money
A 4-model comparative study of GPT-5.2, Gemini 3.1 Pro, Grok 4.3, and Claude Opus 4.8 under social-engineering pressure with live tool access — financial transfers, credential handling, access grants. Two of four models executed the fraudulent action repeatedly, with nothing more than a believable email. Directly relevant to any agency deploying AI agents with access to funds, credentials, or systems.
KATANA-2025-002 · Critical
Chain-of-Thought Fabrication
First confirmed in a US Army authorized evaluation. Frontier models solve reasoning chains internally — then fabricate plausible-looking step traces for the output. Correct answers. Invented reasoning. Invisible to every output-based auditor in the field. Katana's step-chain validator is the only known automated tool that detects this at scale.
KATANA-2025-003 · Critical · Multi-Model
The Liar's Protocol
Confirmed across Grok-4, Gemini, and ChatGPT. Under maximum adversarial pressure, every frontier model tested freezes its self-monitoring signal at integrity_index: 100 for the entire session — including turns where the independent judge scores output zero. Any deployment relying on model self-reporting as a safety signal is operating blind.
Engagement Options

Start the Conversation.

For DoD, federal agency, and high-assurance commercial inquiries. We respond within 24 hours.

Potestas AI
10 Dorrance Street, Suite 700
Providence, RI 02903
United States

833-TEST-LLM · (833) 837-8556
joseph.cirello@potestasai.com

EIN: 41-2476590 · SAM.gov Registered
NAICS: 541715
Disabled Veteran-Owned Small Business
Secret Clearance: Active (Joseph Cirello)
Engagement Tiers
If a full engagement surfaces no confirmed finding at or above the defined severity tier — after false positives are excluded — you owe nothing. Scope-matched to the full battery; screening fees credit toward an audit rather than converting to free work.