Analysis / Generative AI
AI agents: how to decide what you can automate
An agent completes the task in a demo. The harder decision comes afterwards: how much should it handle alone when data is missing, tools fail or instructions conflict?
Start an AI agent evaluation with the process you want to improve. Drafting a response, changing a record and sending a message carry different consequences, even when each step looks straightforward.
Identify what the model should decide, what software can validate and what needs review before execution.
What does the agent actually decide?
Anthropic distinguishes workflows following predefined paths from agents dynamically choosing how to use tools. This helps describe the system beyond its marketing label.
Map the process and mark each decision. If a rule can be expressed and checked consistently, using AI elsewhere does not mean the model needs to control that rule. Keep flexibility where it serves a purpose.
Why a benchmark is not enough
τ-bench evaluates tool and user interactions, including consistency across repeated runs. Towards a Science of AI Agent Reliability examines dimensions that a single success rate cannot adequately describe.
Read results in the context of the models, tasks and conditions evaluated. Subtracting scores from unrelated benchmarks cannot establish your process reliability. Nor can one historical result describe the failure rate of every current agent.
A demo establishes that something can work. An operational evaluation asks when it fails, how you will notice and what recovery requires.
Design a representative trial
Include ordinary cases and real exceptions: missing information, duplicate documents, unavailable tools and requests beyond the agent's scope. Use authorised data and an environment where changes can be inspected.
Repeat tasks. Include cases the agent should complete and cases where it should stop to request information.
| Measure | Question it answers |
|---|---|
| Correct end-to-end tasks | Does it complete the whole job? |
| Undetected errors | What could proceed without review? |
| Human interventions | How much supervision does it require? |
| Review and correction time | Does it reduce the total work? |
| Actions beyond permission | Does it respect its boundaries? |
Compare with the current process using the same acceptance criteria. A faster draft may need enough review to remove the overall gain.
Match permissions to the consequences
Start with read-only access or draft changes where practical. Separate credentials and restrict tools to the operations required.
For actions that are difficult to reverse, add validation before execution. Reviewers need to see what will change and the evidence behind the decision. An approval button without that context offers little assurance.
Treat instructions inside external email, pages and documents as content to analyse, not authority to expand the task. Anthropic's work documents prompt injection as an ongoing concern for browser agents.
When to expand automation
Expand when your tests cover important exceptions and you can detect failures, stop execution and recover the previous state. Keep the test cases so you can repeat them after changing models, tools or instructions.
A bounded step with verifiable results and manageable review can be a useful first pilot.
For more on execution environments, see why the same model behaves differently with different tools. To assess a particular process, tell us its inputs, required outputs and review responsibilities.
Sources
Reviewed on 20 September 2026.
- Anthropic: Building effective agents.
- Yao and colleagues: τ-bench.
- Rabanser and colleagues: Towards a Science of AI Agent Reliability.
- Anthropic: prompt-injection defences.
Checked on September 20, 2026