All notes

Evaluation Inside the Agent Loop

Originally published on LinkedIn

The cheapest high-resolution agent eval may already be inside your execution loop. Just one prediction before each action can turn every run into a labeled dataset.

Poobesh Gowtham ran Claude Code on Opus 5 with a 129-line skill and scored 100.00 on ARC-AGI-3, video games that never explain their rules (for scale, bare Opus 5 held the top spot at 30.2%, GPT-5.6 Sol sits at 7.8%, and eight of the eleven entries are under 1%. Humans finish these games with no instructions)

The mechanism is one rule. Before the agent presses a button, it writes down what the press will do to the grid. cell 22,17=b. The harness refuses any press that arrives without a prediction, so it is a gate and not a line of prompt text asking nicely. That produced 7,627 graded claims, and 443 of them missed. No run exists without the gate, so strictly the harness caused the jump, and the harness is mostly grading machinery.

The pair is what I care about 👀 The outcome metric was pinned at its ceiling with no resolution left to give. Underneath it ran a 5.8% belief-error stream that split hard by mode: single exploratory presses missed 37.1% of the time, planned sequences 2.9%. Nobody labeled those presses as exploration or exploitation. The miss rate found the boundary on its own, which is a live confidence signal you will never get by asking a model how sure it is.

Most of what we call agent evaluation today is final inspection (run the trajectory, score the end, pass or fail). Deming settled the general version of this decades ago: you cannot inspect quality in at the end, and a terminal pass/fail gives you no lever on the process that produced it. A prediction gate is in-process control instead.

The economics are the part I keep pointing people at - the environment was going to return the next state regardless, and that information was already flowing past unused. The claim is what turns a free observation into a labeled one. It produced 7,627 labeled examples with no annotators and no LLM judge in the loop, because the grader is a parser: eight claim forms, closed vocabulary, roughly 10KB of Python, with nothing in it that can drift.

So the work moves! You stop writing better rubrics and start designing a claim grammar that is cheap to state and expensive to hedge. cell 22,17=b cannot be hedged. "The cursor will probably move" cannot be graded.

This needs an environment that returns an observable state delta. Tool calls, migrations, CRM writes, pipeline transforms all have a next frame. Open-ended generation does not. So the question worth asking is what in your system plays the role of the frame.

One line from that repo I keep going back to. Doctrine and harness are co-designed there, meaning every rule the agent is given is one the harness can enforce or grade. Measured that way, a lot of what we ship as agent instructions is not a control at all, it is a wish :)

Back to writing