← RuleReceipt
how it compares

The honest comparison

There are good tools in this space, and most of them are local, free and open-source — same as us. So this page says what each one does better, not just where we fit. If you only read one line: they mostly check a fixed list of signals or block actions in-flight; we read the session afterwards and check it against your own written CLAUDE.md / AGENTS.md rules, quoting the exact line as evidence.

At a glance

Where each tool is strongest, and how we differ. Every claim links to that tool's own docs or repo.

Tool Where it’s genuinely better Overlap & how we differ License · data
RuleReceipt Checks your free-text rules, not a fixed list; quotes the exact transcript line as evidence. Detects & reports after the fact. Only a narrow file/branch guard blocks in-flight. Doesn’t make the model obey. MIT · local, no upload on check
Claude Code hooks Actually preventive — a PreToolUse hook can block an action before it runs. The gate can fail open (per Anthropic’s docs). We catch, after the fact, what it missed. Deep dive → Built-in · local
claude-md-doctor Audits the CLAUDE.md file itself (dead refs, size, unused rules) and proposes a reminder→warn→block ladder. Closest overlap: it also replays rules against transcript history. It leans file-maintenance; we lean per-rule transcript verdict. Deep dive → MIT · local
agent-strace Live control & multi-tool observability — a real-time kill-switch on cost / file / tool-count that stops a run mid-flight. It records and limits; it doesn’t give per-rule followed/broken verdicts against your CLAUDE.md. Its CLI telemetry is on by default (ours isn’t). MIT · anon telemetry on by default
Model Citizen Broader governance & prevention — in-flight hook enforcement, command risk-grading and sandboxing across Claude Code + Codex. Overlaps only in keeping a per-rule ledger. It’s a control-plane; we’re a post-hoc receipt against your own rules. MIT · local, OTLP export off by default
blackbox Cross-session scoring & trends — a weighted score over time and a weekly /retro. Closest in spirit (local, no telemetry, post-hoc Claude Code). But it checks 7 hardcoded signals, not your written rules. Small/early. MIT · local, no telemetry
rulereach Complementary, not competing — a pre-flight linter catching rules that will silently fail to load (broken globs, @import chains, size limits). It answers “will the agent ever load these rules?” We answer “given the session, did it follow them?” Pre-flight vs post-flight. MIT · local
Ask the agent itself Zero setup — nothing to install. Documented LLM self-preference bias makes self-grading unreliable, worst exactly when its own output is wrong. We use the quoted transcript line, not the model’s self-judgment. —

Table scrolls sideways on a phone — the full detail is in the cards below.

The detail, tool by tool

Enough to decide honestly — including the parts where we’re not the right pick.

Claude Code hooks & permissions

Better than us at one thing that matters: a PreToolUse hook can block an action before it runs. That is real prevention, which a post-hoc checker like ours is not. If your only goal is to stop a specific command, start there.

The honest wedge is that the gate can fail open. Anthropic’s own docs say a timed-out hook “doesn’t block the tool call… so don’t count on a stalled hook to act as a gate,” and that a hook “can deny the call, but staying silent doesn’t approve it.” We check every rule after the fact and quote the line — we catch what the gate missed. Full comparison →

claude-md-doctor — the closest overlap

This is the one that overlaps most. Its “session adherence” exam already replays your rules against transcript history with per-rule compliance percentages and quoted evidence — the same idea we built.

Where it does more than us: it also audits the CLAUDE.md file itself — dead references, file size, rules that are never used — and proposes a reminder→warn→block arming ladder. If maintaining a healthy rules file is your problem, it goes further than we do. The fair split: it leans CLAUDE.md maintenance, we lean per-rule transcript verdict. MIT, runs locally. Full comparison →

agent-strace

Better at live control: it’s an observability and replay tool with a real-time kill-switch — cost, file and tool-count limits that stop the agent mid-run, across multiple tools. If you want to cap a runaway session while it’s happening, that’s its job, not ours.

It does not produce per-rule followed/broken verdicts against your CLAUDE.md — that’s the gap we fill. One difference worth knowing: its CLI ships with anonymous telemetry on by default; plain rulereceipt check makes zero network calls. MIT (README-stated).

github.com/JakeSelby/model-citizen (formerly agent-harness)

Model Citizen

Broader and more preventive than us: a full control-plane with in-flight hook enforcement, command risk-grading and sandboxing across Claude Code and Codex. For org-wide governance with live prevention, it does far more than a post-hoc receipt.

We overlap only in keeping a per-rule ledger. Our narrower job — read the finished session, grade it against your own written rules, quote the line — is a different product. MIT (README-stated); local, OTLP export off by default.

blackbox — closest in spirit

Closest to us in philosophy: local, no telemetry, post-hoc Claude Code session compliance. Where it goes further: cross-session scoring and trends, plus a weekly /retro — we focus on a single session’s receipt.

The real difference is what it checks: 7 hardcoded signals (read-before-edit, test-before-commit, destructive commands, correction count, a weighted score) — not your written rules. It can’t grade a rule you wrote in your own words; that’s exactly what we do. It’s small and early (a couple of stars, last pushed mid-2026), so weigh maturity accordingly — for both of us.

rulereach — use it with us, not instead

Not a competitor — a complement. It’s a static pre-flight linter that answers a question we can’t see: will the agent ever load these rules at all? It catches malformed Cursor globs, broken @import chains, size limits and shadowing across Claude, Cursor, Copilot and Codex.

Pre-flight vs post-flight: rulereach checks the rules will reach the model; we check whether the session that followed actually obeyed them. It catches silent load-failures that would otherwise make our “rule followed” meaningless. Run both. MIT (LICENSE-verified).

“Just ask the agent to check itself”

It’s free and needs no setup, and for a quick gut-check it’s fine. But as a reliable check it has a documented problem: LLM self-preference bias (NeurIPS 2024), with a follow-up finding models are least reliable exactly when their own output is the thing that’s wrong — the worst possible time for a self-grade.

Our modest claim: we lean on the quoted transcript line as evidence rather than the model’s self-judgment. That doesn’t make us bias-free — the judgment rules still use a model — but the verdict is anchored to what the transcript actually says, which a self-report inside the same session isn’t.

Two we deliberately left out

Named so you know we looked, not because they earned a row.

There’s a single-purpose “tests-pass” tripwire project (an “actually”-style live check) — a two-day-old repo doing one narrow thing. Worth a mention, not a full comparison. And “rulestack” isn’t a product at all — it’s a dev.to post plus a Gumroad rule-pack for sale — so it isn’t comparable here.

Every tool above is open-source and local, so nobody here is “the only private one” — we’re not, and we won’t say we are. Our two real differences: we check the rules you wrote in plain language, and we quote the exact transcript line as the evidence. If another tool fits your problem better, use it — several pair well with this one.