“How autonomous is the agent?” is becoming a standard product question, and it is the wrong one. Autonomy is not a scalar feature that can be turned from low to high. It is a bundle of permissions exercised across systems with radically different consequences. Reading a repository, opening a pull request, merging to production and rotating a secret may all belong to one workflow, but they should never belong to one undifferentiated authority.
Teams often begin with a chat interface and add tools until the demo feels capable. Each tool appears harmless in isolation. The danger emerges through composition: email supplies untrusted instructions, cloud storage contains sensitive context, a shell can transform files and an issue tracker can trigger humans or automations. A model does not need malicious intent to create damage. It needs only a plausible misunderstanding plus an oversized keyring.
The dominant response is to keep a human “in the loop.” That phrase sounds reassuring while concealing most of the design work. Which loop? At what point? Presented with what evidence? An approval dialog after an agent has already modified a thousand records is ceremonial oversight. Effective control must arrive before the irreversible transition, when the reviewer can still understand the proposed change.
Reversibility is the missing design axis
Every agent action should be classified by blast radius and reversibility. Reading public documentation is broad but low-risk. Drafting a message is reversible until it is sent. Editing a branch is recoverable; rewriting production history is not. Issuing a refund may be appropriate within a limit but dangerous without one. This classification should determine credentials, approval requirements, logging and the environment in which the action runs.
The principle is familiar in security engineering: grant the least authority required for the task. Agent systems make it more important because the principal exercising that authority is probabilistic. The Model Context Protocol architecture, for example, distinguishes hosts, clients and servers and places security boundaries around how context and tools are exposed. A protocol cannot choose a company’s risk tolerance, but explicit boundaries create places where that choice can be enforced.
Reversibility turns those boundaries into a product strategy. Let an agent act freely inside a disposable branch, temporary database or simulated transaction environment. Require a structured proposal before changes cross into durable state. Preserve a before-and-after diff. Attach the evidence used to make the decision. Make rollback an ordinary operation, not an emergency procedure invented after an incident.
Separate planning from authority
A capable agent can be allowed to plan broadly without being allowed to execute broadly. It may inspect a problem, propose a sequence of actions and identify required permissions. A policy layer can then authorize individual steps based on identity, environment, value and confidence. This division is more robust than asking the same model that devised a plan to decide whether its own plan is safe.
Anthropic’s practical guide to building effective agents draws a useful distinction between predefined workflows and agents that dynamically direct their own process. Product teams should extend that distinction to authorization. Flexibility in reasoning does not require flexibility in credentials. The agent may choose how to investigate a failing deployment while still being unable to touch production without an external gate.
This approach also improves usability. Blanket confirmation prompts train people to click “approve” without reading. Good approval design interrupts only at meaningful boundaries and presents a compact decision packet: intended action, affected resources, expected result, uncertainty and rollback plan. The user should not have to reconstruct the agent’s reasoning from a transcript stretching across dozens of tool calls.
Audit logs should explain counterfactuals
Traditional logs answer what happened. Agent audit trails should also help answer what almost happened and why it did not. Record denied tool calls, abandoned plans, policy interventions and the context that caused the agent to change course. These counterfactual traces reveal weak permissions and confusing instructions before they become incidents.
Do not mistake exhaustive logging for accountability, however. A million opaque events merely relocate the problem to the investigator. Logs need stable action identifiers, normalized resources and links between proposals, approvals, executions and reversals. Sensitive reasoning text should not become a new data leak; store the operational evidence necessary to reproduce decisions rather than every token by default.
XioX’s position is straightforward: agent quality should be judged partly by how safely the system fails. A well-designed agent encounters locked doors, bounded accounts and environments that can be reset. It communicates when authority is missing instead of improvising around the restriction. Its designers assume errors will occur and make the cheapest errors the easiest ones to commit.
The race to demonstrate autonomy encourages products to remove friction indiscriminately. Mature systems will do something subtler. They will eliminate friction inside reversible spaces and add precise friction at consequential boundaries. The best agent will not be the one trusted with every key. It will be the one that consistently knows which key it needs—and can proceed safely when it does not have it.
Advertisement