All work

Developer tools · Verification evidence

proof-of-done

Check a coding agent’s verification claims against command results after its latest edits, with a Claude Code stop hook and a transcript audit CLI.

Explore the repository

Verification has to survive the next edit.

An agent can run the tests, change the code again, and still report that everything passes. proof-of-done checks claims such as “tests pass,” “build succeeds,” or “deployed” against the evidence recorded in the session.

The Claude Code plugin runs when an agent or subagent tries to finish its turn. If a recognized claim lacks current supporting evidence, the hook can block the response, explain what is missing, and suggest the command to run.

Follow the edits and command results.

A deterministic detector identifies verification claims. The evidence engine then looks for the latest matching foreground command after the last relevant edit. Failed, empty, interrupted, background, and inconclusive runs do not establish success.

The same engine provides a CLI for auditing past transcripts, exporting findings, and enforcing thresholds in CI. It works locally without calling a model or rerunning commands. Claude Code and agent-trace/v1 inputs are supported; the experimental Codex adapter is for retrospective audits only.

Measure the gate, including its misses.

In the repository’s evaluation of 169 held-out synthetic turns, the default gate blocked 47 of 64 turns containing unsupported claims. It incorrectly blocked one of the other 105 turns. A separate adversarial set met the expected outcome in 34 of 36 cases.

Most held-out errors were wording the detector did not recognize, including some type-check and deployment claims. The published evaluation includes those failures, confidence intervals, and the underlying results.

Current scope.

The checker recognizes English claims and uses shell-command evidence. It cannot see edits made outside the transcript or independently verify an external outcome. Generic “done” and “implemented” statements are outside its scope.

The hook fails open on errors and caps consecutive blocks to avoid trapping a session. Tests exercise the real launcher with fixture transcripts; a live Claude Code smoke run has not yet been completed.