A benchmark for a language model can resemble an exam: present a question, collect an answer, compare it with a reference. An evaluation for a computer-using agent is closer to staging a small world. The evaluator chooses the applications, account state, network conditions, available tools, time limits, and definition of success. Those choices do not merely measure the agent. They shape the behavior being measured.
This distinction matters as agents move from generating text to changing state. A research assistant may search, download, filter, and cite. A workplace agent may update a spreadsheet, send a message, or alter a customer record. Each action changes what the agent encounters next. Two trajectories can reach the same valid result, while two nearly identical trajectories can leave the environment in materially different states.
The influential WebArena paper addressed part of this problem by placing agents inside reproducible websites and grading functional outcomes. OSWorld expanded the idea to real desktop applications and cross-application tasks. These projects made evaluation more realistic, but their deeper lesson is not simply that realistic tasks are harder. It is that an agent score is inseparable from the world in which the score was produced.
Success is not a single bit
Consider an agent asked to reconcile a spreadsheet against invoices. A binary grader might check whether the final totals match. That misses whether the agent silently deleted an unmatched row, overwrote a formula, exposed sensitive data to an unnecessary tool, or arrived at the correct answer through an unreproducible accident. Conversely, a trajectory grader might penalize an unfamiliar but perfectly sound workflow.
Useful evaluations therefore need layers. The outcome layer asks whether the intended state was reached. The constraint layer checks whether prohibited actions occurred. The process layer records retries, reversals, unnecessary tool calls, and requests for clarification. A robustness layer varies details that should not matter: window size, row ordering, filenames, latency, or the location of a button.
No single aggregate score can preserve all of this information. A deployment team should want a profile: completion rate, severe-error rate, intervention frequency, cost, latency, and sensitivity to environmental variation. Collapsing that profile into one leaderboard number is convenient for comparison and poor for engineering.
The environment is executable policy
An evaluation environment contains assumptions about how an agent ought to behave. If the only route to success requires clicking through a warning, the benchmark teaches that warnings are obstacles. If asking the user a question is scored as failure, the benchmark rewards guessing. If a task resets after every run, it hides the operational cost of accumulated mistakes.
This is why benchmark design should be treated as product design. The environment needs explicit affordances for abstention, escalation, and reversible action. Some tasks should reward the agent for stopping when authority is ambiguous. Others should introduce conflicting evidence midway through a workflow and test whether the agent revises its plan. A capable agent is not merely one that continues acting; it is one that recognizes when action has become unjustified.
The grader also needs adversarial testing. Can an agent satisfy the checker while violating the user’s intent? Can it manipulate a visible confirmation field without completing the underlying operation? Can stale state from one trial leak into another? These are not peripheral implementation bugs. They define the validity of the result.
Measure systems, not model folklore
Agent performance depends on the model, prompt, tool descriptions, browser or shell harness, retry policy, context management, and infrastructure. Anthropic’s practical guide to evaluating agents usefully separates code-based, model-based, and human graders. The broader implication is that a benchmark result belongs to a complete system configuration, not to a model name floating free of its scaffolding.
Reports should therefore include the operational details needed to interpret a score: tool APIs, timeout rules, token and step budgets, environment versions, failure recovery, and repeated-run variance. Results from a single run should be viewed with suspicion, especially on long tasks where a minor early error can redirect the entire trajectory.
The next generation of agent evaluation will look less like standardized testing and more like reliability engineering. It will use scenario suites, fault injection, state inspection, and post-incident analysis. It will distinguish a harmless detour from an irreversible mistake. Most importantly, it will test whether an agent knows the boundary of its authority.
The industry wants a clean answer to “How capable is this agent?” The scientifically honest answer is another question: capable under which world, with which tools, under which rules, and at what cost when it is wrong? Better benchmarks will not eliminate that complexity. They will make it legible.
Advertisement