2 articles
The dangerous question is no longer whether generated code looks plausible. Engineering teams need to know what evidence justifies every change and who owns the uncertainty that remains.
Coding-agent evaluations are getting harder, but a larger test suite is not the same as a better measurement. The next useful benchmark will judge recovery, restraint, and maintainability—not merely whether a patch turns the checks green.