[All work](https://boiko.ai/work/)

Agent evaluation · Claim verification

# agent-claimcheck

Check AI agent success claims against tool receipts and state evidence, with calibrated decisions and a review queue for uncertain cases.

[Explore the repository](https://github.com/B0yko/agent-claimcheck)

---

## [Check what actually happened.](https://boiko.ai/work/agent-claimcheck/#check-what-actually-happened)

An agent can report success after an API error, a write to the wrong record, or an operation that never persisted. agent-claimcheck reads agent traces and compares success claims with tool receipts and supplied state probes. Results distinguish verified claims, false successes, and cases that need human review.

The project provides a command-line checker, a Python API, and a local review dashboard. It consumes the agent-trace/v1 format; the caller supplies the traces and environment checks.

## [Rules, learned features, and calibrated abstention.](https://boiko.ai/work/agent-claimcheck/#rules-learned-features-and-calibrated-abstention)

Declarative rule packs check the evidence for each claim. The default offline cascade uses a logistic-regression classifier when rules are inconclusive; an alternative cascade delegates those cases to an OpenAI-compatible LLM judge. Both apply the same decision thresholds after calibration.

Ground-truth labels and source metadata are excluded from detector inputs. Human review decisions can become labelled training examples, while per-trace evidence and recorded evaluations make the decisions inspectable.

## [Measured trade-offs.](https://boiko.ai/work/agent-claimcheck/#measured-trade-offs)

The bundled benchmark contains 300 synthetic traces across booking, CRM, and coding, with 180 for training and 120 for testing. On the test split, the offline cascade decided 90.8% of traces automatically, caught 41 of 48 false successes, incorrectly verified four, and left three for review. It raised no false alarms among 72 genuine successes.

The LLM-judge cascade increased automatic coverage to 96.7% but incorrectly verified seven false successes. The report includes calibration, cost, prompt comparisons, and results grouped by evidence type, so higher coverage is not presented as uniformly better reliability.

[Recorded benchmark](https://github.com/B0yko/agent-claimcheck/blob/63f1ae233ce9605253b53a3dcf607b6d31f00918/results/v0.1.0/bench.md)[Dataset and limitations](https://github.com/B0yko/agent-claimcheck/blob/63f1ae233ce9605253b53a3dcf607b6d31f00918/benchmark/v1/DATASET_CARD.md)

## [What the results establish.](https://boiko.ai/work/agent-claimcheck/#what-the-results-establish)

The benchmark uses simulated tools and templated English traces. Its generator, rule packs, and classifier features share an author; results do not establish performance on real agent runs. Built-in calibrators were fitted on that synthetic distribution and need validation on the intended workload.

Evidence matters: the rules achieved 0.934 AUROC on traces with state probes and 0.750 on receipt-only traces. This is a comparison of benchmark subsets, not an experiment removing probes from identical runs. The tool evaluates the evidence supplied to it; it does not execute agents or independently inspect their live environments.

[Evaluation protocol](https://github.com/B0yko/agent-claimcheck/blob/63f1ae233ce9605253b53a3dcf607b6d31f00918/docs/adr/0005-evaluation-protocol.md)[Trace format and integration scope](https://github.com/B0yko/agent-claimcheck/blob/63f1ae233ce9605253b53a3dcf607b6d31f00918/docs/interop.md)

## [Related work](https://boiko.ai/work/agent-claimcheck/#related-work-heading)

- [booking-truth: evaluating booking outcomes](https://boiko.ai/work/booking-truth/)

- [AI Reliability Layer](https://boiko.ai/work/ai-reliability-layer/)

Source: [https://boiko.ai/work/agent-claimcheck/](https://boiko.ai/work/agent-claimcheck/)
