Teams often begin an AI-agent project by writing a long instruction: describe the role, enumerate prohibitions, demand caution, and tell the model to ask for help when uncertain. Then they connect it to email, a database, a browser, and an internal API. This reverses the normal order of reliable software design. We should not give a probabilistic component broad authority and hope prose turns it into an access-control system.
A production agent needs a smaller room and better doors. The room is the set of states it can observe and change. The doors are typed, permissioned transitions between those states. If an agent drafts a refund recommendation into a review queue, that is one door. If it can issue any refund to any card by composing an HTTP request, that is a hole in the wall.
Anthropic’s practical guide to building effective agents distinguishes predefined workflows from systems that dynamically direct their own actions. That distinction is useful, but many applications need both. Let the model exercise judgment inside a bounded step, while ordinary code owns permissions, invariants, and irreversible transitions. Autonomy should be local, not atmospheric.
Prompts are policy hints, not enforcement
A prompt can express intent: prefer the customer’s current contract, avoid exposing private notes, escalate unusual cases. It cannot guarantee that the model will interpret every novel situation correctly. Enforcement belongs at the tool boundary. A tool for changing a subscription should accept a validated plan identifier and customer identifier, not an arbitrary SQL statement. A payment tool should impose transaction limits independently of anything the model says. A messaging tool should distinguish draft, preview, and send as separate operations.
This architecture improves more than safety. Narrow tools reduce ambiguity for the model. When names, parameters, preconditions, and error responses are crisp, the agent has fewer plausible but wrong actions to choose from. Failures become legible in logs. Tests can target contractual behavior. The same boundary that constrains damage also raises completion rates.
Design every consequential action around four properties: scope, preview, commitment, and reversal. Scope defines what the action may touch. Preview shows the proposed effect using authoritative current state. Commitment occurs through a separate call, sometimes requiring human approval or a short-lived token. Reversal supplies a compensating action or records why none exists. This pattern is mundane in financial and deployment systems; agents make it necessary everywhere.
Measure trajectories, not polished final messages
An agent can produce an excellent summary after taking a terrible route. It may query the wrong account, retry an expensive operation, disclose data to an unnecessary service, or arrive at the correct state accidentally. Final-answer grading misses these defects. Evaluation must inspect the trajectory: chosen tools, arguments, observations, retries, approvals, and state changes. Anthropic’s discussion of agent evaluations makes the central point that multi-turn systems require several kinds of graders because their behavior cannot be reduced to one output string.
Good test suites should include near-miss identities, stale records, conflicting instructions, partial outages, duplicated events, and attempts to inject commands through retrieved content. The objective is not to prove the agent never errs. It is to demonstrate that likely errors terminate safely. A failed search should not become a fabricated fact. A timeout should not trigger an unbounded retry storm. A suspicious document should not acquire authority merely because the agent can read it.
Permission design should also be dynamic. Reading a product catalog may be continuously allowed. Viewing an individual customer record may require task-specific justification. Sending a message or committing a financial change may require explicit approval. The model can request elevation, but it should not grant elevation to itself. Short-lived capabilities make that separation concrete and leave an auditable record of why authority existed.
The human checkpoint needs useful information
“Ask a human” is not a complete control. If the reviewer sees only an approve button and a fluent paragraph, the agent has outsourced its uncertainty without supplying evidence. A proper checkpoint should show the intended action, affected resources, authoritative inputs, relevant policy, and a concise account of unresolved ambiguity. It should make rejection or modification as easy as approval. Otherwise human oversight becomes latency theater.
At XioX, we prefer to begin agent design with an authority map rather than a conversational mockup. List the systems the agent may observe, the state transitions it may propose, the transitions it may commit, and the party capable of reversing each one. Then design tools that embody that map. Only after those boundaries exist should the team tune the prompt and personality.
The ambition of agents should remain high. They can coordinate messy work that rigid workflows handle poorly. But usefulness does not require universal access or uninterrupted autonomy. A capable worker operates inside institutions made of roles, approvals, ledgers, and appeal paths. Software agents deserve equally serious surroundings. Give them room to reason, doors that reveal where they lead, and locks that do not depend on the agent remembering to be careful.
Advertisement