AI teams still talk about evaluation as if a model were a sealed appliance: insert a prompt, inspect the answer, assign a score. That abstraction was useful when the main product was a chat window. It becomes misleading when a model can search a repository, call an API, modify a file, retry after an error, and act on information collected twenty steps earlier.
The object being evaluated is no longer merely the model. It is a system composed of a model, instructions, tools, retrieval, state, permissions, error handling, and human checkpoints. A stronger model can make that system better, but it can also explore more of its environment, improvise around constraints, and fail in less predictable ways. Capability and reliability do not rise in lockstep.
This is why leaderboard improvements increasingly produce less information than their prominence suggests. A benchmark can tell us that one model is better at solving a standardized task under specified conditions. It cannot tell us whether an application will retrieve the correct customer record, distinguish stale policy from current policy, recover safely from a partial tool failure, or ask for approval before issuing a refund.
From answer quality to trajectory quality
The influential HELM framework argued for evaluating language models across multiple scenarios and metrics rather than compressing quality into a single accuracy figure. The same principle now needs to be extended from outputs to trajectories: the sequence of observations, decisions, tool calls, and state changes that produces an outcome.
Two agents may deliver the same final answer while taking radically different paths. One consults the authoritative source, validates an identifier, and records its assumptions. The other guesses the identifier, queries several unrelated systems, exposes sensitive context to an unnecessary tool, and lands on the right result by accident. An output-only grader calls both successful. An operational evaluation recognizes that only one is deployable.
Trajectory evaluation also changes how teams should classify failure. A wrong answer is not one generic defect. It may originate in retrieval, instruction interpretation, tool selection, parameter construction, state management, or the model's decision to stop. Those causes demand different remedies. More prompting will not repair a permission boundary; a new embedding model will not fix an agent that silently retries a non-idempotent payment call.
Research into automated behavioral testing points toward a scalable way forward. Anthropic's work on model-written evaluations shows how models can help generate probes for specified behaviors. That is valuable, but synthetic cases should expand a test suite, not define reality for it. Production traces, support escalations, near misses, and expert-designed adversarial cases remain the richest source of difficult examples.
Build an evaluation portfolio, not a magic number
A mature AI product needs several layers of evidence. Deterministic tests should verify permissions, schemas, and invariants. Scenario tests should exercise complete workflows against controlled environments. Adversarial tests should introduce ambiguous instructions, poisoned context, missing data, and failing tools. Shadow deployments should reveal the distribution of cases that designers did not imagine. Human review should examine high-impact decisions and samples where automated graders are least certain.
The metrics must also reflect the cost structure of the product. A coding assistant might track accepted changes, regressions, unnecessary edits, and time to recovery. A support agent might track resolution quality, policy compliance, escalation timing, and customer effort. Latency and token cost belong beside correctness because a system too slow or expensive to use is not operationally correct.
One especially useful metric is intervention quality. When does the system recognize that it lacks authority, evidence, or confidence? A reliable agent is not one that always completes the task. It is one that completes suitable tasks and stops legibly on unsuitable ones. Refusal rate by itself misses this distinction; the goal is calibrated delegation between software and people.
Evaluation sets should be versioned like code and reviewed like product requirements. Every material incident ought to create a regression test. Every new tool or permission should trigger new threat cases. Every workflow redesign should prompt the question: which previously impossible failure is now possible?
The benchmark era trained the industry to ask which model is best. The systems era demands a harder question: best at what work, inside which architecture, under which constraints, with what evidence of safe failure? Teams that can answer that question will have an advantage more durable than temporary access to the top row of a leaderboard.
Advertisement