All work

Agent evaluation · Booking reliability

booking-truth

Check whether an AI booking agent did what it promised, using calendar-state evaluation, injected failures, and a guarded reference agent.

Explore the repository

Verify the outcome.

A booking confirmation can sound convincing even when the appointment is missing, duplicated, or in the wrong time zone. booking-truth tests appointment-setting agents against a calendar and CRM sandbox, comparing their final state with what the agent told the user.

The harness accepts agents over HTTP, including n8n webhooks. Its 24 scenarios cover booking, rescheduling, cancellation, ambiguous time zones, API failures, repeated messages, and concurrent requests.

Reliability in the implementation.

The reference agent reads calendar writes back before confirming them, uses idempotency keys to handle retries, and restricts bookings to slots returned by a successful availability check. Calendar outages trigger a handoff; verified outcomes enter a persistent queue for CRM updates.

The project includes Cal.com, Google Calendar, and HubSpot adapters, a lightweight chat widget, and per-trial traces. The evaluator is kept independent of the agent’s own claim-checking code.

Measured under injected failures.

In the recorded benchmark, the guarded agent produced no false-success claims across 120 valid trials. The same agent with its guards disabled produced 17 in 120. Each of the 24 scenarios was repeated five times.

These are results from one model in the bundled sandbox. The real-service adapters are implemented against the documented APIs and have not yet been verified against live accounts.