4 articles
Leaderboards move every few weeks and teams are tired of re-evaluating their stack every time a new model tops one. The teams shipping reliably have mostly stopped chasing benchmark scores and started investing in the evaluation harness that tells them how a model performs on their actual task.
Agent benchmarks can tell you whether a system cleared an obstacle course. They cannot tell you whether it will behave sensibly inside your company—unless evaluation becomes part of the product’s architecture.
Coding-agent evaluations are getting harder, but a larger test suite is not the same as a better measurement. The next useful benchmark will judge recovery, restraint, and maintainability—not merely whether a patch turns the checks green.
The industry still talks as if smarter models automatically become better agents. The harder truth is that once models can act, the bottleneck shifts to observing, scoring, and constraining behavior in the messy conditions where real work happens.