Software teams are learning to review code that arrives without a human-scale story of how it was produced. A coding agent may inspect hundreds of files, try several fixes, run a subset of tests, and return a compact patch with a confident summary. The diff is visible. The reasoning that made the diff plausible is scattered across a long interaction trace—or lost entirely.
This creates a peculiar inversion. Generating a patch is becoming cheaper, while establishing confidence in the patch can become more expensive. Teams that respond by demanding more prose from the agent will get longer explanations, not necessarily stronger evidence. The practical answer is to make provenance part of the development artifact.
A plausible patch is not a reviewable change
Traditional code review relies on compressed human context. The author knows why the issue mattered, which approaches failed, which invariant was fragile, and why the final change is narrow. Reviewers can ask questions because there is an accountable mind on the other side of the diff.
An agent can generate a credible narrative after the fact even when that narrative did not guide its actions. Treating this explanation as ground truth confuses fluency with process integrity. The review package should instead contain independently inspectable evidence: the original task, relevant repository state, files read, commands executed, test results, external documentation consulted, and unresolved uncertainty.
The lesson from Anthropic’s guidance on effective agents is that simple, composable workflows often beat elaborate autonomous machinery. The same principle applies to accountability. A short chain of explicit steps is easier to verify than an opaque loop that keeps acting until the output looks finished.
Provenance belongs beside the diff
We propose treating every substantial agent-authored change as a small evidence bundle. It need not expose private chain-of-thought or preserve every token. It should record operational facts that matter to engineering judgment.
- Scope: the issue, acceptance criteria, permissions, and repository revision supplied to the agent.
- Observations: files, symbols, logs, documentation, and runtime behavior the agent actually inspected.
- Actions: edits, commands, migrations, network requests, and generated artifacts.
- Verification: tests run, tests skipped, static checks, reproduction steps, and before-and-after behavior.
- Uncertainty: assumptions that could not be confirmed and areas that still require human judgment.
This is not bureaucratic garnish. It lets reviewers distinguish a locally justified change from a lucky patch. It also lets incident responders reconstruct why a modification passed review months after the original agent session has disappeared.
The SWE-bench paper helped establish real repository issues as a serious test of language-model capability. Production engineering adds requirements that issue resolution alone cannot capture. A patch may satisfy tests while expanding the security boundary, weakening an invariant, adding needless complexity, or depending on behavior that is true only in the evaluation container.
Verification has to be asymmetric
Agents are fast enough to produce more code than humans can carefully inspect. Trying to scale review linearly defeats much of the economic case. Teams need asymmetric verification: cheap mechanisms that reject broad classes of bad changes and focus human attention on the remaining risk.
Repository-specific guardrails are more valuable than generic reminders to “be careful.” Require an agent touching authentication code to run authorization tests. Block dependency additions unless the task explicitly permits them. Compare database plans for query changes. Detect modifications outside the declared scope. Run generated migrations against realistic data volumes. Make the system prove the properties the organization actually cares about.
These controls should be encoded in tools and continuous integration, not left inside prompts. Prompts influence behavior; enforcement defines the boundary.
Real-world productivity evidence also cautions against assuming that more AI use automatically means faster delivery. A METR study of experienced open-source developers found that, in its particular early-2025 setting, participants took longer with AI tools even while perceiving a speedup. The result should not be universalized across tools or teams. Its sharper implication is that subjective smoothness is a poor operational metric. Verification time, rework, review latency, and escaped defects belong in the productivity calculation.
Accountability must follow authority
An autocomplete suggestion and an agent with shell access are not the same product category. As authority increases, the required evidence should increase with it. A system that only drafts a function may need ordinary review. A system that edits configuration, runs migrations, or interacts with deployment infrastructure needs durable logs, constrained credentials, reversible actions, and explicit approval points.
“Human in the loop” is too vague to serve as a control. Which human, shown what evidence, at which irreversible boundary, with how much time? A rushed approval dialog after an agent has already altered dozens of files provides ceremony without oversight.
Good workflow design moves review earlier. The agent first states the intended scope and verification plan. It gathers evidence before editing. It requests elevated authority only when a concrete action requires it. After the change, it reports deviations from the plan. Humans review decisions rather than replaying an entire hidden process.
The audit trail can improve the software
Provenance is often framed as compliance overhead, but its more immediate value is technical. Evidence bundles expose missing tests, ambiguous ownership, undocumented invariants, and repositories that are difficult for any newcomer to navigate. If an agent cannot determine how to validate a change, the problem may be the system’s engineering legibility rather than the model.
This gives teams a better target than maximizing generated lines. Build environments where correct changes are easy to justify and dangerous changes are hard to conceal. The strongest coding agent will not be the one that appears most independent. It will be the one whose work can be challenged, reproduced, and trusted without asking reviewers to take its confidence on faith.
Advertisement