Check what actually happened.
An agent can report success after an API error, a write to the wrong record, or an operation that never persisted. agent-claimcheck reads agent traces and compares success claims with tool receipts and supplied state probes. Results distinguish verified claims, false successes, and cases that need human review.
The project provides a command-line checker, a Python API, and a local review dashboard. It consumes the agent-trace/v1 format; the caller supplies the traces and environment checks.
Rules, learned features, and calibrated abstention.
Declarative rule packs check the evidence for each claim. The default offline cascade uses a logistic-regression classifier when rules are inconclusive; an alternative cascade delegates those cases to an OpenAI-compatible LLM judge. Both apply the same decision thresholds after calibration.
Ground-truth labels and source metadata are excluded from detector inputs. Human review decisions can become labelled training examples, while per-trace evidence and recorded evaluations make the decisions inspectable.
Measured trade-offs.
The bundled benchmark contains 300 synthetic traces across booking, CRM, and coding, with 180 for training and 120 for testing. On the test split, the offline cascade decided 90.8% of traces automatically, caught 41 of 48 false successes, incorrectly verified four, and left three for review. It raised no false alarms among 72 genuine successes.
The LLM-judge cascade increased automatic coverage to 96.7% but incorrectly verified seven false successes. The report includes calibration, cost, prompt comparisons, and results grouped by evidence type, so higher coverage is not presented as uniformly better reliability.
What the results establish.
The benchmark uses simulated tools and templated English traces. Its generator, rule packs, and classifier features share an author; results do not establish performance on real agent runs. Built-in calibrators were fitted on that synthetic distribution and need validation on the intended workload.
Evidence matters: the rules achieved 0.934 AUROC on traces with state probes and 0.750 on receipt-only traces. This is a comparison of benchmark subsets, not an experiment removing probes from identical runs. The tool evaluates the evidence supplied to it; it does not execute agents or independently inspect their live environments.