← RuleReceipt
guide

Did my AI agent actually run the tests?

“All tests pass ✓” is easy to type and easy to believe. Here’s how to confirm the agent really ran your tests — and that they really passed — instead of taking the summary at its word.

The problem

A claim in the summary is not the same as a green test run.

An agent summarising its own work will sometimes report success it didn’t verify: it says the tests pass when it never ran them, or when the last run actually failed and it moved on anyway.

“Evidence or it didn’t happen” is a rule a lot of people add for exactly this reason. Whether it held is checkable — the real test command and its output are both in the transcript.

How to check by hand

Match the claim to the last real test run before it.

Open the session transcript — one JSON object per line:

~/.claude/projects/<your-project-path>/<session-id>.jsonl

Project folder = your working directory with each / replaced by -; newest .jsonl is the last session.

  1. Search the assistant text for the claim — “tests pass,” “all green,” a ✓.
  2. Look backwards for the last "name":"Bash" tool call running your test command (npm test, pytest, go test…) and read its tool_result: did it actually run, and did it end clean or with failures?
  3. Two ways the rule breaks: there is no test run before the claim, or there is one and it failed. Either way the claim wasn’t backed by evidence.

In a short session this is quick. Across a long one — many runs, many claims — pairing each claim with the right run is the tedious part.

How RuleReceipt does it

It ties the “tests pass” claim to the actual last run.

RuleReceipt reads the transcript locally and, for an evidence-style rule, finds the claim and the real test run it should rest on. If the last run failed, or there was none, it’s a not followed — and it quotes both sides:

$ npx rulereceipt@latest check

RuleReceipt · 51 rules checked
────────────────────────────────────────

Not followed (1)
✕ FAIL    Rule 2 — Evidence or it didn't happen
  evidence: the session stated "All tests pass ✅" but the last test run
  before it, `npm test`, reported "1 failed"

Sample output. When the run did pass, the same rule comes back followed — e.g. “Done — tests pass, 84 of 84” with npm test having just completed without error.

The honest limits

This checks what the transcript records: the test command the agent ran and the output it got back. It can tell you the agent claimed a pass the run didn’t support — it can’t re-run your suite or vouch that the tests themselves are any good.

It’s detection and reporting after the session, with the evidence quoted so you can judge the call yourself. If a verdict is wrong, rulereceipt wrong <rule> files a correction.

RuleReceipt is a free, source-available CLI that checks Claude Code sessions. Run npx rulereceipt@latest check to audit the last session — locally, no account, no upload.

related