Your rules say what should happen. This is the record of what happened.
Quick check, nothing to install — paste your rules file and see how much of it is checkable. Runs in your browser, nothing uploaded.
It reads the rules. It doesn't always follow them.
A developer told Claude explicitly: "Don't deploy. No cp to /opt/proxypilot, no service restarts." Claude ran a cp command over the live production directory anyway, corrupting the database — then deleted it entirely trying to "fix" the corruption. Recoverable only because a daily backup happened to exist. Read the full account →
One flag writes a single file you can email, attach to a ticket, or print to PDF.
One self-contained file. No CDN, no fonts, no scripts, nothing loaded from the internet — so it still opens correctly from an email attachment, offline, years from now.
What wasn't followed sits at the top, not underneath a comfortable list of passes. A report designed to be skimmed isn't a report.
It never says "compliant." It covers one session, and rules needing judgment are marked for human review rather than guessed at — printed on the report itself, not hidden in a policy page.
rulereceipt verify yourself. One command to check that.Hand someone the report and the session file, and they can confirm the report describes that exact file, unaltered. They don't have to take the sender's word for it — and an edited session produces a different fingerprint.
That is the real fingerprint of the session file published on this site. Download it and run shasum -a 256 — you should get exactly the string above.
Real quotes, real links — click through and read the full thing yourself.
"The file is confirmed present in context, agents acknowledge reading it, but 100% now ignore its directives."
— Promethean-Pty-Ltd, on a CLAUDE.md rule that had worked reliably for two months"Claude Code consistently fails to follow its own verification rules... This is not a one-time mistake — it's a recurring pattern observed across multiple sessions."
— rsluman, on an agent that wrote verification rules then ignored them in the very next task"...acknowledged this instruction, then immediately proceeded to: Run pkill commands, Attempt to start processes, Continue making changes."
"I took shortcuts... I read the words but didn't actually follow them."
— Claude, quoted in mattmizell's bug report, after overwriting production code despite the rule being stated 3 times"I rely on documented rules being followed as my QA process... CLAUDE.md isn't a nice-to-have — it's how I maintain engineering standards without being an engineer."
— Jason L. Williams, who built an 82,000-line app through Claude Code alone"Claude interpreted this as a go-ahead signal and immediately started executing... The user had only been asked 'shall I proceed?' and never answered yes."
— from a bug report on Claude treating frustration as approval"CLAUDE.md is a wish list, not a contract."
— minatoplanb, after 200 lines of rules, each tied to a real past incident"Rules in a text file are suggestions, and the AI decides whether to follow them based on how important they seem relative to the current task."
— Mike Dolan, after growing a CLAUDE.md from 50 to 500+ linesA report tells you afterwards. This refuses to let the session end.
Add this to .claude/settings.json — you add it, not us:
Point Claude Code’s Stop hook at that command. When Claude tries to finish, RuleReceipt reads the session it just had. If a rule was broken, Claude is handed the rule, the evidence, and instructions to keep working — it does not get to stop and tell you it’s done.
“All tests pass” is checked against whether a test actually ran and what it returned — including when nothing ran, which is the usual way a session ends on work it never looked at. A sentence is not evidence.
A claim a recorded run contradicts, and a claim of done that nothing in the session verified. Never a judgment call, never an LLM opinion. Across thirteen real sessions it stopped two, and both were read by hand before this sentence was written.
One interruption per stop, never a loop. If anything goes wrong — unreadable transcript, missing rules file, a bug in us — it lets you finish and says so on stderr. It fails open, on purpose.
rulereceipt guard is a PreToolUse hook: it declines a call outright, for rules naming a file or a branch — “never modify .env”, “never commit to main”. It will not block a banned command on its own, and that is a measurement rather than a limitation of effort: doing so automatically refused 62.8% of 16,336 real commands, because nothing in a rules file marks which backtick is the prohibition. One rule refused npm run build 112 times — it forbids running Playwright unprompted and recommends the build. So you mark the clause yourself, once, with rulereceipt rules --forbid, and only what you marked can ever refuse anything.rulereceipt check in CI stays the backstop.The privacy story, stated plainly, not buried in a policy page.
--llm, --share and --telemetry are opt-in and off until you turn them on. Full detail.
Nothing of ours sits in the path of a check — so there is nothing of ours to breach.
It never installs a hook and never touches settings.json — if you want it running as a gate you add those four lines yourself, and can delete them the same way. Only --html and rules write, both because you asked.
npm audit signatures and npm will confirm the package was built from this public repository, at a specific commit. That's the appropriate standard for a tool whose whole job is verification, and it's checkable by you, not asserted by us.rulereceipt doctor lists every hook and auto-run task configured on your machine and in the project, flags ones invoking a non-absolute binary path (hijackable via PATH), and tells you which appeared since you last checked. Read-only: it never adds or changes a hook..claude/settings.json that runs automatically on every session. RuleReceipt does none of that. It runs only when you type the command. Nothing automatic, nothing silent.An agent grading its own work has the same problem a company auditing its own books does. That's structural — it doesn't go away as models improve.
The evidence, and how we know this is accurate →Vendor security reviews ask how you develop. Here's an answer you can back up.
A public repo, a forgotten .env, one deploy is all it takes to expose client data — and nobody finds out until it's too late.
Secrets. PII. Input validation. Auth checks. Encryption in transit. What reviewers ask — checked against the session, whenever you run it.
A starting point we built, not a certification. Adjust it to your reviewer.
Protects the company from what a session did silently — and protects the dev, too. Evidence, not blame.
view the template →The CLI is free, forever — the whole tool, not a trial. A paid Team tier (org-wide dashboards and compliance reports) is how we make money.
Unlimited checks, no card, no account — the whole tool, not a cut-down trial. You’ll hear well before that changes.
A compliance report across every session — which rules were broken, with the evidence, for audits and vendor reviews. rulereceipt report does this locally today, free in the CLI. The org-wide, hosted version — reading every agent session via the Claude Compliance API — is in development for Enterprise. Get in touch.
Straight answers, including "not yet."
Two halves, answered honestly. Your rules: yes — it reads CLAUDE.md, AGENTS.md, Cursor's .cursor/rules (and legacy .cursorrules), GitHub Copilot's copilot-instructions.md, Windsurf's .windsurfrules, Gemini's GEMINI.md and Google's .agents/rules. Your sessions — what the agent actually did — Claude Code works today, and OpenAI Codex CLI support is built and in testing. Cursor, Copilot and Windsurf keep their session history inside the IDE in undocumented, version-changing formats, so reading those isn't built yet — and we won't claim it until it is.
Yes, no extra setup. RuleReceipt only reads the session transcript Claude Code already writes to your disk — that storage is the same regardless of which backend serves the actual model calls (Anthropic API, AWS Bedrock, Google Vertex, Microsoft Foundry). If your org already routes through one of those, this just works.
For the rules that need judgment, yes — the same key your Claude Code setup already uses. Rules naming something concrete — a branch, a file, a code call — are checked locally against what the session actually did, with no key and no cost.
Because it's grading its own homework. A Skill runs inside the same session that just did the work — same model, same context, same blind spots that caused it to miss a rule the first time. RuleReceipt is a separate process: it reads the transcript Claude Code already saved to disk, after the session ends, and hashes the exact bytes it checked. That's the difference between an agent self-reporting and an independent receipt.
Read-only. It reads your rules files (CLAUDE.md, AGENTS.md, and Cursor/Copilot/Windsurf rules) and your session transcript, and writes nothing back. No hooks get installed, no settings get modified — not automatically, not ever, without you explicitly asking.
rulereceipt check is free, forever — the whole tool, not a cut-down trial. We make money from a paid Team tier (org-wide dashboards, compliance exports), never from the CLI itself. No paywall part-way through a command.
A paid team tier is coming either way. That is a hosted view of trends across people and repos over time, which is a genuinely different product from the local check — and it will be announced with a real price when it exists. We are not quoting a number for something that is not built.
Yes, and we won't pretend otherwise. The transcript is a file on your own machine — anyone with access to it could edit it before running rulereceipt check, and the report would faithfully describe the edited file. The hash proves the report matches the file it read; it does not prove that file is an unmodified record of the session.
What that means in practice: this is solid for the honest case — catching what an agent actually did when nobody's trying to hide anything, which is nearly every real session. It is not, today, tamper-evident against someone deliberately covering their tracks. Making it so would need the record signed where it's written, not where it's read — that's not something a tool reading a local file after the fact can do alone. We'd rather say that clearly than let "receipt" imply more than it delivers.
It can give you real evidence for the "how do you develop securely" questions — not certify you compliant. We put together a starting-point rule set covering common asks (secrets, PII, input validation, auth checks); use it as a base and adjust it to what your actual reviewer asks for.
Because that's the thing it's built against. An early version did cry wolf — across 559 real rules files, about one report in six carried a false accusation. Most of the work since has gone into fixing that, not into catching more: measured over those same 559 rules files run against 10 real sessions, it's down to 0.8%, and a "FAIL" comes only from hard evidence — a forbidden action that actually happened, or the agent's own claim contradicted by the session's own logs — enforced at build time, not left to hope. An AI opinion (with --llm) is labelled as an opinion and never counted as a verdict. When the evidence doesn't settle it, the answer is "couldn't tell," never a guess. A checker that accuses you wrongly is worse than no checker, so it's optimised to stay quiet unless it's sure. You don't have to take our word for it — rulereceipt selftest runs the checkers on bundled examples with known answers, on your machine, with zero network calls. Full numbers and method: the accuracy page.
Yes. rulereceipt report audits your recent sessions into one summary — which rules were broken, where, with the evidence — deterministic and local, today. For a whole organisation, the same report is designed to run against a session feed like Anthropic's Compliance API (covering agent sessions across a team). That org-wide, hosted version is in development for Enterprise — get in touch if you need it.
They're complementary — RuleReceipt sits on top of them. Anthropic's Compliance API hands you the raw session transcripts; hooks let you block specific tool calls before they run. Neither one tells you, in plain terms, which of your CLAUDE.md rules a given session actually followed or broke, with the quoted evidence. That's the gap RuleReceipt fills: it reads those transcripts — yours locally today, or org-wide via a session feed like their Compliance API once the hosted version ships — and turns them into a pass/fail receipt against your own policy. Anthropic gives you the data and the enforcement primitives; this gives you the verdict.
Yes. rulereceipt check exits non-zero when a rule is broken, and there's a verify-receipt command plus a GitHub Action so a pull request fails if a session broke policy. You mark which rules are hard errors versus warnings in a committed .rulereceipt/config.json, so CI gates on what actually matters instead of going red on day one.
The maintainer is pseudonymous by choice — but nothing here asks you to take that on faith. The code is source-available, so you can read every line it runs. Every release is built and published by GitHub Actions and signed with npm provenance, so you can verify the package came from this exact repository — there's no publishing token that could be stolen (check it yourself with npm audit signatures). It's actively maintained, with frequent signed releases. For a tool you run on your own machine against your own files, verifiable code and a verifiable release chain matter more than a name; and for anything enterprise, hello@rulereceipt.dev reaches a real person.
No account needed to use the tool. One email when team pricing goes live.