AI benchmarks once worked like surveying markers: fixed reference points that let researchers compare progress across time. That bargain is breaking down. Popular evaluations now circulate through papers, repositories, tutorials, synthetic-data pipelines, and human feedback. Models may encounter benchmark questions, close variants, solution strategies, or the cultural residue of the test long before evaluation day.
The familiar response is to search for cleaner data. That helps, but it treats contamination as a housekeeping failure rather than a structural condition. A benchmark influential enough to guide research will eventually become part of the environment it measures. The more important the test, the harder it is to keep outside the model’s world.
This does not make evaluation futile. It means the field must stop treating a benchmark as a permanent ruler.
A score compresses away the failure surface
The original HELM framework made a crucial move by evaluating models across scenarios and multiple dimensions, including robustness, calibration, fairness, and efficiency. Its enduring lesson is not that one suite can be sufficiently comprehensive. It is that capability has a shape, and a single aggregate score hides that shape.
A model can improve on a public reasoning set while becoming more brittle when instructions conflict. It can produce better code patches while relying on repository conventions it has effectively memorized. It can gain factual accuracy yet become harder to calibrate because fluent wrong answers look increasingly convincing. None of these behaviors is captured adequately by a victory-lap number on a leaderboard.
Evaluation should therefore begin with a failure hypothesis. What would make this system unsafe, uneconomic, or simply annoying in the intended setting? For a support assistant, it might be inventing refund authority. For a coding agent, it might be changing unrelated files or passing visible tests through an implementation that violates an unstated invariant. For a research tool, it might be producing citations that exist but do not support the surrounding sentence.
The test follows from the hypothesis. Starting with a benchmark and working backward encourages teams to optimize what is convenient to measure.
Freshness is a system property
Software engineering provides a clear illustration. SWE-bench improved on toy programming tests by grounding tasks in real issues and repositories. That was a meaningful advance. But any fixed collection of public issues acquires a half-life: patches become searchable, repository patterns become familiar, and agent scaffolds adapt to the evaluation harness.
Later work on SWE-bench-Live points toward a better model: continuously updated tasks, executable environments, and attention to contamination resistance. The important word is “live.” Freshness cannot be a release-day property. It has to be maintained through collection, versioning, isolation, and retirement.
We think serious evaluation programs will increasingly resemble observability platforms. They will ingest new cases, preserve traces, segment results by failure mode, and alert teams when behavior shifts. A test set will be one component inside that system, not the system itself.
- Public canonical tasks will remain useful for comparability and debugging.
- Private rotating tasks will test whether gains transfer beyond familiar examples.
- Adversarial variants will perturb wording, tools, context, permissions, and environmental state.
- Production traces will reveal failures that benchmark authors did not anticipate.
- Human review will examine whether a nominally successful answer used an acceptable path.
Measure trajectories, not only answers
Agentic systems make endpoint scoring especially inadequate. Two agents can reach the same answer while creating radically different risk. One inspects relevant files, forms a hypothesis, runs targeted tests, and makes a narrow patch. Another searches broadly, rewrites several modules, retries until tests turn green, and leaves behind an accidental dependency. A binary pass records equivalence where an engineering team sees opposite outcomes.
Evaluation should capture the trajectory: tool calls, files touched, permissions requested, discarded hypotheses, test selection, token and compute cost, and recovery after failure. The OpenAI Evals repository helped normalize the idea that teams should build evaluations around their own use cases. The next step is to treat traces as first-class evidence rather than diagnostic exhaust.
This changes what “better” means. A stronger system is not merely one that completes more tasks. It reaches correct outcomes with fewer unjustified actions, detects uncertainty earlier, and degrades predictably when the environment changes. Efficiency and restraint are capability metrics.
The evaluation team should be allowed to surprise the product team
Many internal eval programs are structurally unable to discover bad news. Product teams define the expected workflows, construct representative cases, and decide when the result is good enough. That arrangement produces useful regression tests, but it also bakes the product’s assumptions into the measurement.
A credible evaluation function needs some independence. It should be able to introduce tasks from outside the happy path, withhold portions of the distribution, and report disaggregated failures even when the headline average rises. It should also publish the boundaries of its claims: languages not tested, tools not available, user groups underrepresented, and environmental conditions held constant.
The deepest shift is conceptual. Benchmarks are not trophies awarded to models. They are instruments maintained by institutions. Instruments drift, environments change, and measurement affects the thing being measured. The labs and companies that recognize this will learn faster than those chasing the cleanest-looking score. They will also be more willing to say the sentence every trustworthy evaluation should make possible: the model improved here, failed there, and we know why.
Advertisement