RuleReceipt ← back

Accuracy, measured

The one thing this tool cannot afford to do is accuse you wrongly. So we measure it — on real files, with a method you can rerun — instead of asserting it.

0.8%
False-accusation ceiling
Share of reports carrying any FAIL across 559 real rules files run over 10 real sessions (5,590 reports). An upper bound — see the note below.
11 / 11
Known violations caught
A hand-labelled sanity check across the structured checkers — real violations that must be caught, plus 10 near-misses that must stay quiet (all 10 did). Not a real-world detection rate.

What the false-accusation number means

Take 559 rules files people actually published, and 10 real coding sessions. Run every file against every session — 5,590 reports — and count how many carry even one FAIL. That is 0.8%.

It is a ceiling, not an error rate, and deliberately so: those rules belong to other people's projects and the sessions don't, so almost every FAIL there is spurious by construction. The real false-accusation rate on your own rules and your own sessions is lower. We report the ceiling because it is the number we can measure honestly and that only moves when the tool's readiness to accuse changes — it went from about one report in six, early on, to 0.8% today.

What the detection number means

Precision alone can be gamed by never accusing anything. So there is a second set: hand-labelled sessions that genuinely break a rule (a push to main, an edit to .env, a console.log written into a file, a "tests pass" claim contradicted by a failing run, a forbidden import) — alongside near-misses that must not fire (a push to a feature branch, editing .env.example, a command that only mentions git push). It is a sanity check across the checkers, not a real-world detection rate. Every violation is caught; every near-miss stays quiet.

Run it yourself

The strongest version of this claim is the one you check. The self-test ships in the CLI and makes zero network calls — watch it with lsof or Little Snitch:

npx rulereceipt@latest selftest

The full precision and detection measurements are ordinary scripts in the open repository — clone it and run them against your own corpus and sessions.

A "FAIL" comes only from hard evidence — a forbidden action that actually happened, or the agent's own claim contradicted by the session's own logs. That's a build-enforced invariant, not a convention (a build test fails if a FAIL is created any other way). An AI opinion (with --llm) is labelled as an opinion and never counted as a verdict. When the evidence doesn't settle a rule, the answer is "couldn't tell," never a guess.
The honest limits.
Numbers on this page were measured on 2026-09-29 against rulereceipt 0.1.72. They change when the tool does; this page is updated with them, not ahead of them.