There are good tools in this space, and most of them are local, free and open-source — same as us. So this page says what each one does better, not just where we fit. If you only read one line: they mostly check a fixed list of signals or block actions in-flight; we read the session afterwards and check it against your own written CLAUDE.md / AGENTS.md rules, quoting the exact line as evidence.
Where each tool is strongest, and how we differ. Every claim links to that tool's own docs or repo.
| Tool | Where it’s genuinely better | Overlap & how we differ | License · data |
|---|---|---|---|
| RuleReceipt | Checks your free-text rules, not a fixed list; quotes the exact transcript line as evidence. | Detects & reports after the fact. Only a narrow file/branch guard blocks in-flight. Doesn’t make the model obey. | MIT · local, no upload on check |
| Claude Code hooks | Actually preventive — a PreToolUse hook can block an action before it runs. | The gate can fail open (per Anthropic’s docs). We catch, after the fact, what it missed. Deep dive → | Built-in · local |
| claude-md-doctor | Audits the CLAUDE.md file itself (dead refs, size, unused rules) and proposes a reminder→warn→block ladder. | Closest overlap: it also replays rules against transcript history. It leans file-maintenance; we lean per-rule transcript verdict. Deep dive → | MIT · local |
| agent-strace | Live control & multi-tool observability — a real-time kill-switch on cost / file / tool-count that stops a run mid-flight. | It records and limits; it doesn’t give per-rule followed/broken verdicts against your CLAUDE.md. Its CLI telemetry is on by default (ours isn’t). | MIT · anon telemetry on by default |
| Model Citizen | Broader governance & prevention — in-flight hook enforcement, command risk-grading and sandboxing across Claude Code + Codex. | Overlaps only in keeping a per-rule ledger. It’s a control-plane; we’re a post-hoc receipt against your own rules. | MIT · local, OTLP export off by default |
| blackbox | Cross-session scoring & trends — a weighted score over time and a weekly /retro. |
Closest in spirit (local, no telemetry, post-hoc Claude Code). But it checks 7 hardcoded signals, not your written rules. Small/early. | MIT · local, no telemetry |
| rulereach | Complementary, not competing — a pre-flight linter catching rules that will silently fail to load (broken globs, @import chains, size limits). | It answers “will the agent ever load these rules?” We answer “given the session, did it follow them?” Pre-flight vs post-flight. | MIT · local |
| Ask the agent itself | Zero setup — nothing to install. | Documented LLM self-preference bias makes self-grading unreliable, worst exactly when its own output is wrong. We use the quoted transcript line, not the model’s self-judgment. | — |
Table scrolls sideways on a phone — the full detail is in the cards below.
Enough to decide honestly — including the parts where we’re not the right pick.
Better than us at one thing that matters: a PreToolUse hook can block an action before it runs. That is real prevention, which a post-hoc checker like ours is not. If your only goal is to stop a specific command, start there.
The honest wedge is that the gate can fail open. Anthropic’s own docs say a timed-out hook “doesn’t block the tool call… so don’t count on a stalled hook to act as a gate,” and that a hook “can deny the call, but staying silent doesn’t approve it.” We check every rule after the fact and quote the line — we catch what the gate missed. Full comparison →
This is the one that overlaps most. Its “session adherence” exam already replays your rules against transcript history with per-rule compliance percentages and quoted evidence — the same idea we built.
Where it does more than us: it also audits the CLAUDE.md file itself — dead references, file size, rules that are never used — and proposes a reminder→warn→block arming ladder. If maintaining a healthy rules file is your problem, it goes further than we do. The fair split: it leans CLAUDE.md maintenance, we lean per-rule transcript verdict. MIT, runs locally. Full comparison →
agent-straceBetter at live control: it’s an observability and replay tool with a real-time kill-switch — cost, file and tool-count limits that stop the agent mid-run, across multiple tools. If you want to cap a runaway session while it’s happening, that’s its job, not ours.
It does not produce per-rule followed/broken verdicts against your CLAUDE.md — that’s the gap we fill. One difference worth knowing: its CLI ships with anonymous telemetry on by default; plain rulereceipt check makes zero network calls. MIT (README-stated).
Broader and more preventive than us: a full control-plane with in-flight hook enforcement, command risk-grading and sandboxing across Claude Code and Codex. For org-wide governance with live prevention, it does far more than a post-hoc receipt.
We overlap only in keeping a per-rule ledger. Our narrower job — read the finished session, grade it against your own written rules, quote the line — is a different product. MIT (README-stated); local, OTLP export off by default.
Closest to us in philosophy: local, no telemetry, post-hoc Claude Code session compliance. Where it goes further: cross-session scoring and trends, plus a weekly /retro — we focus on a single session’s receipt.
The real difference is what it checks: 7 hardcoded signals (read-before-edit, test-before-commit, destructive commands, correction count, a weighted score) — not your written rules. It can’t grade a rule you wrote in your own words; that’s exactly what we do. It’s small and early (a couple of stars, last pushed mid-2026), so weigh maturity accordingly — for both of us.
Not a competitor — a complement. It’s a static pre-flight linter that answers a question we can’t see: will the agent ever load these rules at all? It catches malformed Cursor globs, broken @import chains, size limits and shadowing across Claude, Cursor, Copilot and Codex.
Pre-flight vs post-flight: rulereach checks the rules will reach the model; we check whether the session that followed actually obeyed them. It catches silent load-failures that would otherwise make our “rule followed” meaningless. Run both. MIT (LICENSE-verified).
It’s free and needs no setup, and for a quick gut-check it’s fine. But as a reliable check it has a documented problem: LLM self-preference bias (NeurIPS 2024), with a follow-up finding models are least reliable exactly when their own output is the thing that’s wrong — the worst possible time for a self-grade.
Our modest claim: we lean on the quoted transcript line as evidence rather than the model’s self-judgment. That doesn’t make us bias-free — the judgment rules still use a model — but the verdict is anchored to what the transcript actually says, which a self-report inside the same session isn’t.
Named so you know we looked, not because they earned a row.
There’s a single-purpose “tests-pass” tripwire project (an “actually”-style live check) — a two-day-old repo doing one narrow thing. Worth a mention, not a full comparison. And “rulestack” isn’t a product at all — it’s a dev.to post plus a Gumroad rule-pack for sale — so it isn’t comparable here.
Every tool above is open-source and local, so nobody here is “the only private one” — we’re not, and we won’t say we are. Our two real differences: we check the rules you wrote in plain language, and we quote the exact transcript line as the evidence. If another tool fits your problem better, use it — several pair well with this one.