2 articles
Leaderboards move every few weeks and teams are tired of re-evaluating their stack every time a new model tops one. The teams shipping reliably have mostly stopped chasing benchmark scores and started investing in the evaluation harness that tells them how a model performs on their actual task.
The industry still talks as if smarter models automatically become better agents. The harder truth is that once models can act, the bottleneck shifts to observing, scoring, and constraining behavior in the messy conditions where real work happens.