Calling an AI system a “digital employee” is an effective sales tactic and a poor design specification. Employees use judgment shaped by history, relationships, norms, and consequences. A model receives context, emits outputs, and perhaps invokes tools. Treating those things as equivalent encourages businesses to automate a role when they should be redesigning a flow of work.
The distinction is not semantic. Roles are bundles of loosely connected responsibilities: investigate a complaint, calm a customer, interpret policy, update records, spot a recurring issue, and tell a manager when the policy itself is wrong. Software performs better when that bundle is decomposed into bounded transitions with visible inputs and outputs.
Instead of giving an agent a job title, give it a queue, a budget, and an escalation rule. The queue defines what work may enter. The budget limits time, tool calls, money, and retries. The escalation rule identifies the conditions under which uncertainty returns to a person. These three elements turn an impressive demo into an operable service.
The workflow is the product
Most organizations begin with the model: choose a provider, build a chat interface, connect a few tools, and look for tasks. That sequence overvalues conversational fluency. The better starting point is an actual work item and the state changes it must pass through.
Consider invoice exceptions. A robust system does not need to “be an accounts-payable specialist.” It needs to extract fields, compare them with a purchase order, classify the discrepancy, request missing evidence, recommend an action within policy, and route unusual cases for approval. Each transition can be observed and tested. The model may handle several transitions, but it does not own the process.
Anthropic’s guide to building effective agents draws a useful distinction between workflows, where code determines the path, and agents, where a model dynamically directs its own process. That distinction should be a design decision, not a maturity ladder. More autonomy is not inherently more advanced. If the task has stable stages and clear rules, a workflow is usually easier to test, cheaper to run, and safer to change.
Reserve open-ended agency for the sections that genuinely require it: searching unfamiliar evidence, forming a plan for a novel case, or choosing among tools when the route cannot be known in advance. Surround those sections with deterministic software. Autonomy should appear as a bounded pocket inside a system, not as a fog spread over the entire operation.
Make state legible
Chat transcripts are a poor operational database. They mix instructions, evidence, speculation, and decisions in one stream. When something goes wrong, a team must reconstruct what the system believed and which action changed the world.
Production workflows need explicit state: received, validated, awaiting evidence, proposed, approved, executed, or escalated. They also need structured records of sources, tool results, policy versions, and approvals. The model can contribute to state transitions, but authoritative state should live outside its prose.
This separation makes recovery possible. If a downstream service fails, the system can retry one idempotent step rather than replaying an entire conversation. If policy changes, pending cases can be identified. If an auditor asks why an action occurred, the answer can point to evidence and authority rather than to a long, probabilistic transcript.
It also clarifies what humans are for. “Human in the loop” often means placing a person at the end of a noisy conveyor belt to approve whatever the model produced. That is not oversight; it is alert fatigue with a friendly name. Human attention should be allocated according to risk and information value.
- Low-risk, reversible actions can run automatically with sampling.
- Ambiguous cases should request a targeted judgment, not a review of the entire history.
- High-impact or irreversible actions should require explicit authority before execution.
- Novel failures should feed process redesign, not merely produce one-off corrections.
Measure prevented work, not generated output
AI pilots commonly report how many drafts, summaries, or classifications the system generated. Output volume is easy to count and weakly connected to value. An agent can generate thousands of artifacts while creating a new review burden.
Better measures follow the work: elapsed time to resolution, touches per case, rework, exception rate, unauthorized actions prevented, and the share of cases completed without hiding uncertainty. Cost should include human verification and incident handling, not just token spend.
The NIST AI Risk Management Framework offers a durable vocabulary for governing and measuring AI risk. Its practical relevance is that governance cannot be bolted on after automation. Authority, monitoring, and accountability have to be represented in the workflow itself.
Practitioners should also follow independent experimenters who expose the messy edges of real tools. Simon Willison’s writing, for example, consistently treats model behavior as something to probe with working software rather than accept from product language. That habit—test the actual boundary—is more useful than debating whether a system philosophically qualifies as an agent.
The deepest opportunity is not replacing a person-shaped box on an organization chart. It is removing avoidable waiting, copying, searching, and coordination while preserving the moments where responsibility matters. That may change roles, but role reduction is an outcome of better process design, not a suitable starting objective.
Companies that anthropomorphize their systems will keep discovering that apparent initiative is not accountability. Companies that engineer queues, budgets, state transitions, and escalation paths will build something less theatrical and more valuable: operations that can explain what happened, recover when assumptions fail, and improve from every exception.
Advertisement