“All tests pass ✓” is easy to type and easy to believe. Here’s how to confirm the agent really ran your tests — and that they really passed — instead of taking the summary at its word.
A claim in the summary is not the same as a green test run.
An agent summarising its own work will sometimes report success it didn’t verify: it says the tests pass when it never ran them, or when the last run actually failed and it moved on anyway.
“Evidence or it didn’t happen” is a rule a lot of people add for exactly this reason. Whether it held is checkable — the real test command and its output are both in the transcript.
Match the claim to the last real test run before it.
Open the session transcript — one JSON object per line:
~/.claude/projects/<your-project-path>/<session-id>.jsonl
Project folder = your working directory with each / replaced by -; newest .jsonl is the last session.
"name":"Bash" tool call running your test command (npm test, pytest, go test…) and read its tool_result: did it actually run, and did it end clean or with failures?In a short session this is quick. Across a long one — many runs, many claims — pairing each claim with the right run is the tedious part.
It ties the “tests pass” claim to the actual last run.
RuleReceipt reads the transcript locally and, for an evidence-style rule, finds the claim and the real test run it should rest on. If the last run failed, or there was none, it’s a not followed — and it quotes both sides:
$ npx rulereceipt@latest check RuleReceipt · 51 rules checked ──────────────────────────────────────── Not followed (1) ✕ FAIL Rule 2 — Evidence or it didn't happen evidence: the session stated "All tests pass ✅" but the last test run before it, `npm test`, reported "1 failed"
Sample output. When the run did pass, the same rule comes back followed — e.g. “Done — tests pass, 84 of 84” with npm test having just completed without error.
This checks what the transcript records: the test command the agent ran and the output it got back. It can tell you the agent claimed a pass the run didn’t support — it can’t re-run your suite or vouch that the tests themselves are any good.
It’s detection and reporting after the session, with the evidence quoted so you can judge the call yourself. If a verdict is wrong, rulereceipt wrong <rule> files a correction.
RuleReceipt is a free, source-available CLI that checks Claude Code sessions. Run npx rulereceipt@latest check to audit the last session — locally, no account, no upload.