Imagine you're in a conference room. Opposing counsel, a regulator, or your board asks: "Show us evidence that you verified the AI output before you acted on it." What do you hand them?
These questions are no longer hypothetical. A recent federal sanctions case highlighted the gap between having an AI governance policy and proving you followed it. The court didn't just want to know if the law firm had rules about AI use. It wanted records showing verification actually happened. When the firm couldn't produce them, it got sanctioned under Rule 11, a requirement that's been in place since 1937.
This case is important because it shows what "evidence-ready AI governance" really means. Your team needs to move beyond Policy Gap Analysis and ask: If we had to prove compliance tomorrow, what would we produce?
The Current State of AI Governance
Most organizations have some version of an AI policy. Many include terms like "human oversight" or "verification of AI outputs." These are assertions, not evidence. The Western District of Tennessee case involving Reaves Law Firm illustrated this gap clearly.
The firm filed court pleadings that defendants alleged contained AI hallucinations: arguments unsupported by cited cases, quotes that didn't exist. The court issued a show-cause order requiring Reaves to confirm whether the cases existed and, critically, to "identify what steps, if any, it had taken prior to citing to the cases to verify their existence."
Reaves produced an internal email titled "Mandatory Ethical AI Training & Reporting Protocols for All Staff." But the firm couldn't demonstrate the email was distributed firm-wide or that training occurred. It blamed a departed general counsel and mentioned "restructured" processes. The court sanctioned the firm anyway.
The lesson isn't that the firm lacked a policy. It's that the firm couldn't produce records showing the policy was followed.
Is Your AI Governance Evidence-Ready?
Ask yourself: Can you produce contemporaneous documentation of AI oversight decisions? Not a policy statement saying "we review AI outputs," but actual records showing who reviewed what, when, and what judgment they exercised.
Evidence-ready governance requires four capabilities: Define, Record, Own, and Guard.
Define means specifying which decisions require human review. Not "we use AI responsibly," but "these specific outputs require verification by this role before this action." In the Reaves case, filing court pleadings is obviously consequential. But the firm never demonstrated a clear answer about AI's role in producing those pleadings.
Record means creating evidence of what the AI produced and how the human processed it. What was reviewed? What standard was applied? What conclusion was reached? Reaves couldn't produce verification records when the court asked for them.
Own means assigning accountability to a named person. Reaves tried to blame a departed general counsel, but the court held the firm and the signing attorney accountable. When you can't point to who verified what, accountability becomes diffuse and unenforceable.
Guard means monitoring your governance over time and escalating when signals appear. Reaves received allegations of hallucinations in the first pleading, then filed two more with the same defects. A warning arrived, and no one stopped to review or correct the process.
Why "Human in the Loop" Isn't Enough
"Human in the loop" is a design principle, not evidence of compliance. When someone tests that assertion, you need records proving the human was actually in the loop for specific decisions.
Consider what you'd need to produce if an auditor asked: "Show me evidence that a qualified person reviewed the AI output before you relied on it in this decision." You'd need a log entry, a review record, or an approval trail that shows who reviewed what output, when, and what they verified. A policy document saying "we do human review" doesn't answer that question.
The Reaves court didn't create a new AI-specific obligation. Rule 11 has required verification of legal citations for decades. AI just made it easier to delegate a task that carries a duty and harder to prove appropriate oversight after the fact.
What Does a Defensible Verification Record Look Like?
It depends on the decision's risk and complexity, but at minimum you need: what AI output was produced, who reviewed it, what they verified, and what action they authorized based on that review.
For high-stakes decisions, you might need more: the prompt or input that generated the output, the standard or criteria applied during review, any discrepancies identified, and how those discrepancies were resolved. If your process involves sampling rather than reviewing every output, document the sampling methodology and criteria.
The record should be contemporaneous. Reconstructing verification steps from memory months later, especially after personnel changes, is unreliable and unconvincing. Reaves discovered this when it tried to explain its processes after its general counsel had left.
Does a Small Team Need This Level of Documentation?
If you're making consequential decisions based on AI output, yes. The size of your organization doesn't change the nature of the duty.
Start with high-risk decisions where AI influences outcomes that could harm people, violate regulations, or expose the organization to liability. Define what verification means for those decisions and create a simple record when it happens. A spreadsheet tracking reviews is better than nothing. A workflow system with approval steps is better still.
The Reaves case involved court filings, but the principle applies anywhere AI output influences decisions with external consequences: regulatory submissions, financial disclosures, hiring decisions, credit determinations, medical recommendations.
Auditing AI Governance Amidst Technological Change
Focus on the governance infrastructure, not the technology. Your audit scope should cover whether the organization can answer four questions for each AI use case:
- Is there a defined standard for when and how AI output requires human review?
- Are there records showing that review actually occurred?
- Is there a named person accountable for each review decision?
- Is there a process to detect and respond when the review process fails?
You're not auditing the AI model itself. You're auditing whether the organization has operationalized its policy claims with evidence-generating processes.
Starting Point for Policy-Heavy AI Governance
Inventory your high-risk AI use cases. For each one, ask: If we had to prove compliance tomorrow, what would we produce?
If the answer is "our policy document," you have work to do. Build the Define-Record-Own-Guard capabilities for your highest-risk applications first. Create templates for verification records. Assign ownership. Establish escalation triggers.
Then test it. Simulate a show-cause scenario: "Prove you verified this AI output before acting on it." If you can produce contemporaneous records showing who reviewed what and what judgment they exercised, you're evidence-ready. If you can't, you know where the gap is.
The question isn't whether your AI policy is well-written. It's whether you can prove you followed it.




