AI teams have inherited a comforting ritual from conventional machine learning: choose a benchmark, run the model, compare the number. The ritual feels scientific because the output has decimals. It is also becoming dangerously incomplete. A general-purpose model is not a fixed classifier operating inside a tidy distribution. It is a probabilistic component wrapped in prompts, retrieval, tools, policies, and application state. What users experience is that entire system, yet companies still make purchasing and deployment decisions from model-level scores.
This gap matters because benchmarks answer narrow questions. They can tell us whether one model solved more items from a particular set under a particular harness. They rarely tell us whether an insurance assistant will recognize an ambiguous exclusion, whether a coding agent will recover from a malformed dependency, or whether a support bot will escalate when the customer is distressed. Those are not exotic edge cases. They are the work.
Public scores decay faster than teams admit
A benchmark begins losing information as soon as it becomes influential. Model builders optimize against it, examples and explanations circulate online, and the test distribution becomes familiar. Research on benchmark data contamination explains why exposure can make apparent generalization difficult to separate from memorization. Even without literal leakage, saturation compresses meaningful differences near the top of a leaderboard.
The common response is to add another benchmark. That produces a wider spreadsheet, not necessarily better evidence. Ten static tests may share the same blind spots: short tasks, clean inputs, obvious stopping conditions, and answers that can be graded automatically. Real deployments involve incomplete instructions, conflicting evidence, delayed consequences, and users who change their minds halfway through.
A better unit of evaluation is the capability claim. “This agent can fix software bugs” is too broad. “This agent can diagnose and patch a reproducible defect in this repository class, without changing public behavior elsewhere, under this tool budget” can be tested. The narrower statement may sound less impressive, but it can support an engineering decision.
Measure trajectories, not just final answers
For an agentic system, the route to an answer is part of the result. Two agents may produce identical patches while taking radically different paths. One inspected the failing test, formed a plausible hypothesis, changed one module, and reran relevant checks. The other searched indiscriminately, modified six files, consumed ten times the tokens, and happened to pass. A binary success metric calls them equal. An operator should not.
Evaluation therefore needs to capture tool calls, intermediate decisions, retries, cost, latency, reversibility, and the quality of escalation. METR’s work on task-completion time horizons is valuable partly because it reframes capability around sustained completion of work, rather than isolated puzzle solving. The broader lesson is that duration and dependency structure change what competence means.
This does not require grading hidden reasoning. In many applications, observable behavior is enough. Did the agent inspect the governing policy before issuing a refund? Did it preserve an audit trail? Did it notice that two sources disagreed? Did it stop when permissions were missing? These are testable properties of a trajectory.
Build an evaluation portfolio like a risk portfolio
No test format can carry the whole burden. XioX recommends a portfolio with four layers:
Stable regression tests catch failures in behavior the product already promises. They should run often and change slowly.
Fresh challenge sets use newly written or procedurally varied cases to reduce overfitting and contamination.
Adversarial scenarios target costly failure modes: prompt injection, misleading evidence, excessive authority, and confident fabrication.
Production sampling examines what real users actually attempt, with privacy controls and human review appropriate to the domain.
The portfolio should be weighted by consequence, not convenience. If a wrong answer can move money or expose private data, a rare catastrophic failure deserves more attention than a frequent formatting defect. Aggregate accuracy conceals this asymmetry. A system that is excellent on routine cases and reckless on exceptional ones can look strong in an average.
Evaluation sets should also have owners and expiration dates. An unowned test suite becomes a museum of last year’s concerns. Product, domain, security, and engineering teams should regularly add cases from incidents, near misses, and changed workflows. Some tests should remain hidden from the people tuning prompts and agents. Otherwise the organization recreates the public-leaderboard problem internally.
The goal is not a perfect score
A living evaluation system does more than approve model releases. It reveals where the surrounding product needs stronger constraints. A failure may call for a better model, but it may instead call for narrower permissions, a deterministic calculator, improved retrieval, or a mandatory human checkpoint. Treating every miss as a prompting problem is how brittle systems acquire elaborate prompts and unchanged risk.
The most mature teams will stop asking, “Which model wins?” and ask, “What evidence justifies this system performing this task under these conditions?” That question is less marketable and far more useful. It turns evaluation from a leaderboard ceremony into institutional memory: a growing record of what the product claims, where it has failed, and what must remain true when any component changes.
The benchmark is not obsolete. It is simply one instrument on the bench. The mistake is confusing the instrument with the measurement program.
Advertisement