Teams often begin an agent project by assigning it a job title. The system becomes a research assistant, sales representative, support specialist, or engineering teammate. This feels concrete, but it is weak architecture. A title describes a social expectation; it does not define what software may read, change, spend, disclose, or send.
The dangerous sentence in an agent specification is not “act as an operations manager.” It is “use the available tools to complete the task.” Available to whom, under which identity, against what scope, with which limits, and with what recovery path? If those questions are unanswered, the prompt is being asked to perform the work of an authorization system.
Prompts cannot carry that burden. They are probabilistic instructions interpreted alongside untrusted documents, messages, websites, and tool outputs. The more useful an agent becomes, the more adversarial or simply confusing material it encounters. A polite sentence telling it not to reveal secrets is not equivalent to preventing its process from reading those secrets in the first place.
Anthropic’s guide to building effective agents draws a valuable distinction between fixed workflows and systems that direct their own tool use. That autonomy can handle ambiguity, but it also moves critical decisions into the agent loop. The correct response is not to eliminate autonomy. It is to give autonomy a carefully engineered perimeter.
Authority should be smaller than capability
A model may be capable of drafting a refund, querying a customer record, changing an account, and sending an email. It does not follow that one agent session should possess all four powers simultaneously. Capability describes what the system can reason about. Authority describes what this particular invocation is allowed to do.
Good deployments make authority narrower than capability. A support agent can inspect an order but receive only a purpose-built refund tool with a transaction ceiling. A research agent can read approved sources but cannot publish externally. A coding agent can modify a working branch but cannot alter repository protections or production secrets. These are not model-behavior preferences. They are enforced properties of the surrounding system.
Each tool should therefore be treated as a grant of power. Its interface ought to encode the smallest useful action, rather than exposing a general-purpose escape hatch. “Issue refund for this order within policy” is safer than unrestricted database access. “Create draft message” is safer than “send email.” Narrow tools also improve reliability because they reduce the action space the model must navigate.
Approval prompts are not a complete safety model
Human confirmation is useful when an action is consequential and unusual. It becomes theater when users face dozens of context-poor prompts. People habituate, approve reflexively, or cannot realistically reconstruct the agent’s prior steps. The approval dialog then transfers liability without transferring understanding.
Use approvals at authority boundaries, not as punctuation after every tool call. Reading a project file inside an authorized workspace may require no interruption. Sending its contents outside the organization should. Editing a reversible draft may be routine. Executing a payment, deleting a record, or widening access should require a fresh decision supported by a concise preview of effects.
This approach aligns with the broader risk-management logic in the NIST generative AI profile: risks must be governed, mapped, measured, and managed across the system lifecycle. For an agent, that means the unit of governance cannot stop at the model. Identity, tools, data stores, logs, user interfaces, and incident procedures are part of the deployed system.
Design an authority envelope
Before writing the agent’s persona, write its authority envelope. At minimum, specify:
- Identity: whether the agent acts as itself, the requesting user, or a dedicated service account;
- Resources: the exact repositories, records, folders, or accounts it may access;
- Actions: which operations are read-only, reversible, approval-gated, or prohibited;
- Budgets: limits on money, tokens, time, recipients, and action frequency;
- Destinations: where information may be transmitted and in what form;
- Lifetime: when credentials and delegated authority expire;
- Evidence: what must be logged so a reviewer can reconstruct decisions;
- Recovery: how changes are rolled back and who owns escalation.
This envelope should be enforced below the prompt layer. Use scoped credentials, transactional APIs, sandboxed execution, allowlisted destinations, immutable audit records, and default-deny policies. The agent can reason inside those constraints without being trusted to preserve them.
Prompt injection makes this separation especially urgent. Guidance from the OWASP GenAI Security Project treats prompt injection and excessive agency as system-level concerns, not defects solved by clever phrasing. Any content an agent reads can attempt to redirect it. The robust assumption is that instructions and data will sometimes be confused, so unauthorized actions must remain unavailable even when reasoning fails.
Permissions can improve the product
Security constraints are often framed as friction, but explicit authority makes agents easier to understand. Users can delegate confidently when they know the boundary. An interface that says “may read these three folders and create drafts, but cannot send or delete” is more useful than a vague promise that the assistant will be careful.
Boundaries also improve evaluation. Teams can test whether every attempted action stayed within scope, whether approval was requested at the right moment, and whether expired authority was rejected. Those outcomes are observable. “Behaved like a responsible colleague” is not.
The strongest agent products will not be those that imitate employees most convincingly. They will be those that make delegation precise. Give the model enough room to plan, investigate, and adapt—but surround that freedom with enforceable limits designed for software, not social intuition. A job title belongs in an org chart. An agent needs a permission boundary.
Advertisement