A benchmark score is irresistible because it compresses uncertainty into a number. An agent receives a task, produces an answer, and either passes or fails. Model A beats Model B; the purchasing decision appears settled. But an agent is not a function that maps one prompt to one response. It is a process that interprets a goal, chooses tools, changes external state, reacts to observations, and decides when to stop. Evaluating only its final answer is like testing a database by inspecting one exported spreadsheet.
This distinction matters as agents move from demonstrations into software engineering, research, support, and operations. The original SWE-bench paper improved the field by testing models on real repository issues rather than toy code completion. Yet even a realistic task corpus cannot tell a product team why its particular agent succeeds. The result may depend on repository setup, tool descriptions, search strategy, retry policy, context management, or sheer luck. A headline score conceals the machinery that users actually encounter.
Successful outcomes can hide defective behavior
Suppose a coding agent fixes a bug and passes the test suite. The run looks successful. Its trajectory, however, reveals that it read a credential file it did not need, rewrote an unrelated configuration, attempted the same failed command six times, and consumed twenty times the expected tokens. A binary grader awards a pass. An engineering organization should see four incidents.
The inverse is also true. An agent may take a disciplined path, identify the correct root cause, and fail because a package mirror briefly disappears. Calling that a capability failure corrupts the evaluation. Outcome, process, and environment must be recorded separately. Otherwise teams optimize prompts to compensate for infrastructure noise or blame models for broken tools.
Anthropic's guide to agent evaluations usefully emphasizes that multi-turn systems require multiple grader types. We would push the idea further: an agent evaluation should resemble an observability system. It should capture the decisions that preceded the output, the state changed along the way, and the conditions under which a human would have wanted the run interrupted.
Measure the trajectory as a typed object
A useful evaluation record needs more structure than a transcript. At minimum, it should identify the initial state, permitted actions, tool calls, state mutations, recoverable errors, irreversible actions, cost, latency, stopping reason, and final state. Each field answers a different product question. Did the model understand the task? Did the tool contract make the right action discoverable? Did a retry improve the result? Did the agent stop because it finished, because it exhausted a budget, or because it mistakenly believed it had finished?
This produces a practical set of evaluation layers:
- Task success: was the requested result achieved under realistic conditions?
- Path quality: did the agent take relevant, economical, and policy-compliant steps?
- State integrity: were unrelated files, records, permissions, and external systems left untouched?
- Recovery behavior: did the agent distinguish transient failures from evidence that its plan was wrong?
- Calibration: did it request help when uncertainty or consequence justified escalation?
Not every action deserves equal weight. An unnecessary read is wasteful; an unnecessary payment is dangerous. Evaluation therefore needs consequence-aware grading. The same behavioral pattern can be harmless in a sandbox and unacceptable in production. This is one reason simple, composable agent patterns are easier to ship responsibly: their control points remain legible enough to test.
Build evals from operating history, not imagination
Teams commonly create evaluation sets before launch, then treat them as a permanent exam. That reverses the useful relationship between product and evaluation. Production should continuously supply new cases: confusing requests, ambiguous permissions, unusual tool responses, costly loops, and near misses caught by operators. After sanitization, these become regression scenarios. The evaluation suite becomes institutional memory for the system.
Freshness matters because static benchmarks eventually reward familiarity. The move toward live, newly collected tasks in SWE-bench-Live points in the right direction. Product teams need the same discipline internally. Hold out recent cases, rotate environments, perturb tool output, and test semantically equivalent instructions. If a tiny change in phrasing or state causes a collapse, the system has not learned a robust procedure; it has found a narrow groove.
Statistical humility belongs here too. Agents are stochastic, and difficult tasks are often sparse. A few passes do not establish reliability. Run repeated trials, report distributions rather than a single average, and inspect whether failures cluster around a particular tool or decision. The unit of analysis is not merely the model. It is the model-tool-policy-environment combination.
Evals are part of the architecture
The strongest evaluation systems change how agents are built. If state integrity is graded, tools acquire dry-run modes and scoped permissions. If recovery is graded, errors become machine-readable rather than buried in prose. If escalation quality is graded, the product needs an explicit handoff state instead of an apologetic chat message. Measurement exposes missing architecture.
This is the more demanding view of agent progress. Better models matter, but reliability will not arrive as a side effect of a higher public score. It will come from systems that make behavior observable, consequences testable, and failure reusable. The benchmark can tell us whether an agent crossed the finish line. The product evaluation must tell us what it broke, what it learned, and whether we should let it run again.
Advertisement