An agent grading its own work has the same problem a company auditing its own books does. That is structural — it does not go away as models improve.
Three independent sources, each pointing at the same gap.
Their paper Trustworthy agents in practice splits agent security into four layers — model, harness, tools, environment.
“A well-trained model can still be exploited through a poorly configured harness, an overly permissive tool, or an exposed environment.”
Three of those four layers belong to whoever deploys the agent, not to the model vendor. They describe what is needed as “the kind of infrastructure no single company can build alone.”
From a survey of organisations running AI agents. That last figure is the whole problem in a sentence: you cannot review what you cannot distinguish.
The Model AI Governance Framework for Agentic AI — the first of its kind — asks for two things:
We'd rather be straight about the limits of that: it is voluntary guidance, not enforcement, and the EU AI Act's high-risk obligations were pushed to December 2027, not this year. The direction is clear even where the dates aren't — and we're not going to manufacture a deadline out of it.
Because we got it wrong first, in public, and then measured the fix.
We ran this against a real session and it reported ten rule violations. Every one was wrong. Rather than patch them one by one, we found the cause:
A text match can never prove someone did something.
Searching a session for a banned command finds it whether the agent ran it, looked it up, or mentioned it in a note. So we rebuilt it to read what actually happened, and a text match can no longer report a failure at all. Where it isn't certain, it says so instead of guessing.
The whole thing is written up, including the wrong turn we took while fixing it: what we got wrong.
CLAUDE.md, AGENTS.md, .cursorrules, Copilot instructions, Windsurf and Gemini files — from PyTorch, Kubernetes, Elasticsearch and hundreds of other projects.
These figures change when the classifier is corrected, and the earlier values stay published rather than being edited away — see how they have moved, and why.
When that corpus grew roughly fourteen-fold, into formats the classifier had never been tuned on, its behaviour barely moved. That stability is the actual argument: directives are a property of language, not of file format.
And it is tested in the other direction too — sessions built to deliberately break a rule, to confirm it catches them, not merely that it stays quiet on well-behaved ones.
The figures above are dated on purpose — they come from specific 2026 studies and they will age. The structural argument is what we would stand behind regardless: an independent record beats a self-report, whatever the current numbers say.