The most consequential document in an AI product is often not its system prompt, architecture diagram, or model card. It is the evaluation suite. Evals decide which behavior the team notices, which regressions block a release, and which failures remain invisible until customers find them. Yet many teams still treat evaluation as a final exam administered after the product has effectively been designed.
That framing is backward. A useful evaluation is an executable product requirement. It translates a sentence such as “the assistant should help analysts review contracts” into observable decisions: which clauses it must identify, when it must quote evidence, what counts as an unacceptable omission, and when it should decline to answer. Writing those decisions down exposes ambiguities that a polished demo can conceal.
General intelligence is not the unit of deployment
Public benchmarks are valuable because they create shared reference points. Stanford’s HELM project, for example, made the case for evaluating language models across multiple scenarios and metrics rather than compressing capability into one score. The Chatbot Arena research showed how pairwise human preferences can reveal differences that static question sets miss.
But no public benchmark knows what your customer is trying to accomplish. A model can improve on broad reasoning tests while becoming less dependable inside a claims workflow because it now produces more persuasive unsupported explanations. It can win preference comparisons while violating a house style that distinguishes facts from assumptions. “Better model” and “better component for this system” are different claims.
The relevant unit is not abstract intelligence. It is a particular model, equipped with particular tools and instructions, operating on a particular distribution of work. Change the retrieval index, document parser, tool permissions, or interface and the system has changed even if its model identifier has not. The evaluation boundary must surround the deployed behavior, not merely the API call at its center.
An eval suite contains a theory of harm
Every test set encodes priorities. If it contains many easy successes and few costly edge cases, it says average fluency matters more than tail risk. If graders reward answer similarity, it says conformity matters more than justified variation. If every prompt is clean and self-contained, it says the messy state of production is somebody else’s problem.
Strong teams make that theory explicit. They begin with a failure inventory assembled from domain experts, support conversations, incident reviews, and adversarial testing. They separate failures that are merely awkward from those that are expensive, irreversible, or difficult to detect. A fabricated restaurant recommendation and an invented contractual obligation are both hallucinations, but treating them as equivalent produces a dangerously bland metric.
This leads to weighted evaluation. The scorecard for a medical scheduling assistant might assign severe penalties to changing a dosage instruction, moderate penalties to losing an appointment constraint, and small penalties to stylistic repetition. The weights are not universal scientific constants. They are accountable product judgments, which is precisely why they should be visible.
Measure abstention and recovery, not only answers
Many evaluations assume the system must answer every prompt. Production systems have more options. They can ask a clarifying question, retrieve another source, request approval, route the case to a specialist, or state that available evidence is insufficient. Those behaviors often determine whether an uncertain model becomes a dependable service.
An eval should therefore test trajectories. Did the system choose an appropriate tool? Did it preserve source context through several steps? Did it notice contradictory evidence? After a failed call, did it recover without repeating an external action? The final sentence may look correct even when the path was reckless.
That is also why automated model graders need calibration. They are useful for coverage and iteration speed, but they inherit preferences and blind spots of their own. A practical regime mixes deterministic checks, model-based grading, and periodic human review. Disagreements between graders are not noise to discard; they identify cases where the specification is underspecified.
The test set should age with the product
A static benchmark becomes ceremonial. Users adapt, source data changes, attackers discover new seams, and the team adds features that create unexpected interactions. Each meaningful production failure should become a durable regression case, with sensitive details removed. Each major workflow change should trigger a review of whether the test distribution still resembles actual use.
The healthiest evaluation suites are versioned alongside the application. Their results are segmented by task and risk rather than hidden inside one average. Release reports explain which capabilities improved, which weakened, and which remain uncertain. A team may knowingly accept a regression, but it should have to name the tradeoff.
This practice changes product conversations. Instead of asking whether the newest model is smarter, teams ask whether the system now fails inside the tolerances their users can bear. That question is less glamorous and far more useful. An eval suite is not a scoreboard hanging outside the product. It is the product’s operational definition of good judgment.
Advertisement