AI-assisted programming has made code cheap to produce and strangely expensive to trust. A change can arrive with tidy abstractions, passing surface-level tests, and a persuasive explanation. None of those properties establishes that the change belongs in production.
Many teams respond by keeping their existing pull-request process and adding a human glance at the end. That is a weak control. Review practices evolved when code production was constrained by developer time and authors carried substantial context from problem to implementation. A coding agent can generate a broad patch quickly, while the reviewer inherits neither the author’s struggle nor a reliable account of what was considered and rejected.
The remedy is not to ban generated code or demand that humans inspect every line with equal intensity. It is to build a chain of custody for changes: a record connecting the original intent, repository evidence, implementation decisions, validation results, and accountable human approval.
A diff is the end of the story, not the beginning
Conventional review centers on the diff because the author can answer questions about its origin. With an agent, the useful artifact is larger. What instruction did it receive? Which files and documentation did it inspect? What assumptions did it make about callers, data, and deployment? Which tests failed during iteration? What remained unverified?
This does not mean storing every token of an agent transcript. Raw transcripts are noisy and can create a false sense of auditability. Teams need a compact evidence bundle: task statement, constraints, touched interfaces, tests added or run, tool outputs that support key claims, and explicit uncertainty. The bundle should be generated during the work, not reconstructed after a reviewer finds something suspicious.
SWE-bench helped establish a valuable standard for coding agents: proposed fixes should be judged by executable tests against real repositories. Yet production engineering asks a wider question than whether a patch resolves a benchmark issue. A test suite reflects the behavior its authors anticipated. Generated code can preserve those expectations while violating performance budgets, observability conventions, security boundaries, or undocumented operational knowledge.
Review the risk surface, not the volume
Line count is a poor proxy for danger. A one-line authorization change can be more consequential than a generated test fixture spanning hundreds of lines. Review depth should follow the risk surface: privilege boundaries, irreversible data changes, public APIs, concurrency, money movement, personal data, and failure recovery.
Teams can encode this idea in workflow. Low-risk changes may proceed when targeted tests, static checks, and ownership rules agree. High-risk changes should require a design note, stronger test evidence, and review by someone responsible for the affected domain. An agent may help assemble that evidence, but it cannot approve its own assumptions.
This is also why broad prompts such as “clean up this service” are operationally hazardous. They allow the agent to choose both the objective and the solution. Better tasks specify the permitted scope, invariants that must remain true, observable success conditions, and actions that require confirmation. Precision at delegation time reduces review ambiguity later.
Passing tests are evidence, not absolution
Executable validation is the strongest advantage software work has over many other knowledge tasks, but it can be misused. If an agent writes both the implementation and the only new tests, the two may share the same misunderstanding. Existing tests can also reward compatibility with accidental behavior.
A stronger pattern separates claims. Contract tests check externally visible behavior. Regression tests reproduce the reported failure. Property or fuzz tests explore cases the patch author did not handpick. Static analysis examines whole classes of defects. Staging and telemetry reveal interactions that local execution cannot. The appropriate mixture depends on the system, but the principle is stable: independent evidence is more valuable than repeated evidence from one reasoning path.
Anthropic’s discussion of agent evaluation distinguishes outcome checks from assessments of the trajectory. That distinction belongs inside engineering workflows too. Did the final tests pass? Good. Did the agent disable a check, silently narrow the task, or repeatedly attempt an unsafe operation before arriving there? The path can reveal risk that the final patch conceals.
Ownership cannot be automated away
The most corrosive failure mode is social rather than technical: nobody feels like the author. The requester supplied a prompt, the agent produced the patch, and the reviewer approved what appeared on screen. When the change fails, responsibility dissolves across the workflow.
Every merged change needs a human owner who can explain why its evidence is sufficient. That person need not have typed the code. Modern engineering already accepts libraries, generated clients, migrations, and compiler output that no reviewer examines character by character. Accountability comes from understanding provenance, constraints, and validation—not from keystrokes.
This reframes the productivity conversation. The valuable metric is not generated lines or completed tickets. It is safely accepted change per unit of expert attention. An agent that produces twice as much code but consumes three times the review effort is not accelerating the system. An agent that creates smaller patches, exposes uncertainty, and supplies targeted evidence may be far more useful even if it appears less autonomous.
AI coding tools will keep getting better at producing plausible implementations. That makes disciplined provenance more necessary, not less. The mature team will not ask whether a patch feels human-written. It will ask what claim the patch makes, what independent evidence supports that claim, what uncertainty remains, and which person is willing to own the answer.
Advertisement