Model evaluations arrive with the visual language of settled science: decimal points, ranked tables, confidence-inducing labels. A model scores 73, another scores 71, and the industry behaves as if somebody placed both systems on a calibrated scale. But most AI benchmarks are not scales. They are assemblies of prompts, graders, tools, sampling settings, infrastructure, and undocumented judgment calls. Change one component and the apparent capability can move.
This does not make benchmarks useless. It makes them measurement systems whose uncertainty has been systematically hidden. The field’s fixation on leaderboard position encourages teams to report an output while ignoring the machinery that produced it. That habit was merely sloppy when evaluations involved static multiple-choice questions. It becomes dangerous when the object under test is an agent operating a browser, terminal, codebase, or simulated workplace.
The harness is part of the capability
Consider software-engineering evaluations. The original SWE-bench paper made a valuable move away from toy code completion toward real repository issues. Yet an agent’s result on such a benchmark depends on far more than the underlying model. The system prompt matters. So do the available commands, repository setup, retry policy, context management, test visibility, token budget, and rules for declaring completion.
Calling all of that “scaffolding” understates the point. The harness is not packaging around intelligence; it is part of the system being measured. A strong model inside a confused interface can fail. A weaker model supplied with a careful search loop, concise tools, and useful feedback can look dramatically better. The benchmark score therefore supports a claim about a particular model-system pairing under particular conditions—not about an abstract model floating free of implementation.
This is familiar territory in experimental science. A temperature reading without the instrument, calibration state, placement, and error tolerance is incomplete. AI evaluation should adopt the same intellectual posture. Every serious result needs an uncertainty budget: which factors could have moved the outcome, which were controlled, and which remain unknown?
Three uncertainties that leaderboards conceal
The first is sampling uncertainty. Agent tasks are often expensive, so teams run surprisingly few trials. But one run can hinge on a tool timeout, an early mistaken branch, or a lucky recovery. Pass-at-one is useful only if “one” is understood as a draw from a distribution. Reporting repeated runs on a meaningful subset can reveal whether a system is dependable or merely occasionally brilliant.
The second is environment uncertainty. Package versions drift. Websites change. Tests can be flaky. Network access alters what the agent can discover. Even machine load can affect long-running tool interactions. Organizations such as METR, which studies the ability of advanced systems to complete extended tasks, are pushing evaluation toward more realistic work. Realism, however, introduces more moving parts. That makes environment capture and replay more important, not less.
The third is construct uncertainty: whether the benchmark measures the capability named in the headline. A coding benchmark may partly measure familiarity with popular repositories. A factuality test may reward conservative refusal. A “reasoning” task may be solvable through recognizable templates. The remedy is not to hunt for one contamination-proof benchmark. It is to state the claim narrowly, design adversarial variants, and test whether performance survives changes that should be irrelevant.
Evaluations should behave like maintained infrastructure
A benchmark is often treated as a publication: release the dataset, create a leaderboard, and move on. Production evaluation should instead resemble observability infrastructure. It needs versioning, provenance, incident review, access controls, and owners. When a score changes, the team should be able to determine whether the model improved, the harness changed, the grader drifted, or the test distribution shifted.
Frameworks such as Stanford’s Holistic Evaluation of Language Models point in the right direction by making scenarios and metrics more explicit. The deeper shift is organizational. Evaluation engineers should have authority comparable to platform engineers, because their work defines what the organization believes about its systems.
At XioX, we think the most useful evaluation artifact is not a single aggregate number. It is a capability map with boundaries: the system succeeds reliably here, fails predictably there, and becomes unstable under these conditions. That map can guide deployment, routing, interface design, and human review. A leaderboard position usually cannot.
The industry does not need fewer benchmarks. It needs fewer unqualified claims. A score becomes meaningful when its test conditions are reproducible, its variance is visible, and its interpretation is specific enough to be wrong. Until then, many benchmark tables are not measurements of intelligence. They are screenshots of complicated experiments whose controls are just outside the frame.
Advertisement