4 articles
Reasoning models do not merely answer tests; they interact with them. Evaluation must therefore measure how systems spend effort, use tools, recover from errors, and behave under changing constraints.
Inference-time computation is turning a single model into a family of systems with different costs, latencies, and capabilities. Evaluating them requires measuring a curve, not publishing one triumphant score.
AI evaluation is drifting toward theater: cleaner leaderboards, weaker understanding. The next serious wave of benchmarks will focus less on whether a model got the final answer and more on how it behaved while getting there.
The AI field has become too comfortable mistaking leaderboard movement for genuine understanding. The harder question is no longer which model wins a benchmark, but whether the benchmark still describes the work we care about.