← RuleReceipt
the case for an independent record

Why an outside check

An agent grading its own work has the same problem a company auditing its own books does. That is structural — it does not go away as models improve.

The argument isn't ours

Three independent sources, each pointing at the same gap.

Anthropic · April 2026

The vendor makes the architectural case themselves

Their paper Trustworthy agents in practice splits agent security into four layers — model, harness, tools, environment.

“A well-trained model can still be exploited through a poorly configured harness, an overly permissive tool, or an exposed environment.”

Three of those four layers belong to whoever deploys the agent, not to the model vendor. They describe what is needed as “the kind of infrastructure no single company can build alone.”

Cloud Security Alliance · April 2026

The gap is measured, not assumed

53%had agents exceed intended permissions
8%said it never happens
2 in 3can't tell agent actions from human ones

From a survey of organisations running AI agents. That last figure is the whole problem in a sentence: you cannot review what you cannot distinguish.

Singapore IMDA · January 2026

Where the rules are heading

The Model AI Governance Framework for Agentic AI — the first of its kind — asks for two things:

  • A verifiable identity per agent.
  • An audit trail of which agent acted under whose authorization.

We'd rather be straight about the limits of that: it is voluntary guidance, not enforcement, and the EU AI Act's high-risk obligations were pushed to December 2027, not this year. The direction is clear even where the dates aren't — and we're not going to manufacture a deadline out of it.

How we know this tool is accurate

Because we got it wrong first, in public, and then measured the fix.

Every failure it first reported was a false alarm

We ran this against a real session and it reported ten rule violations. Every one was wrong. Rather than patch them one by one, we found the cause:

A text match can never prove someone did something.

Searching a session for a banned command finds it whether the agent ran it, looked it up, or mentioned it in a note. So we rebuilt it to read what actually happened, and a text match can no longer report a failure at all. Where it isn't certain, it says so instead of guessing.

The whole thing is written up, including the wrong turn we took while fixing it: what we got wrong.

Tested against real rules files, not our own examples

559public rules files
21,986items parsed
8,101real rules
0crashes

CLAUDE.md, AGENTS.md, .cursorrules, Copilot instructions, Windsurf and Gemini files — from PyTorch, Kubernetes, Elasticsearch and hundreds of other projects.

These figures change when the classifier is corrected, and the earlier values stay published rather than being edited away — see how they have moved, and why.

When that corpus grew roughly fourteen-fold, into formats the classifier had never been tuned on, its behaviour barely moved. That stability is the actual argument: directives are a property of language, not of file format.

And it is tested in the other direction too — sessions built to deliberately break a rule, to confirm it catches them, not merely that it stays quiet on well-behaved ones.

The figures above are dated on purpose — they come from specific 2026 studies and they will age. The structural argument is what we would stand behind regardless: an independent record beats a self-report, whatever the current numbers say.