Skip to content

Analysis / Generative AI

AI agents: how to decide what you can automate

An agent completes the task in a demo. The harder decision comes afterwards: how much should it handle alone when data is missing, tools fail or instructions conflict?

Contrast 3DPublished Updated 3 min read

Ceramic robot arm passing a card through a violet glass gate between two document trays
AI-generated conceptual illustration. Not a screenshot or a photograph.

Start an AI agent evaluation with the process you want to improve. Drafting a response, changing a record and sending a message carry different consequences, even when each step looks straightforward.

Identify what the model should decide, what software can validate and what needs review before execution.

What does the agent actually decide?

Anthropic distinguishes workflows following predefined paths from agents dynamically choosing how to use tools. This helps describe the system beyond its marketing label.

Map the process and mark each decision. If a rule can be expressed and checked consistently, using AI elsewhere does not mean the model needs to control that rule. Keep flexibility where it serves a purpose.

Why a benchmark is not enough

τ-bench evaluates tool and user interactions, including consistency across repeated runs. Towards a Science of AI Agent Reliability examines dimensions that a single success rate cannot adequately describe.

Read results in the context of the models, tasks and conditions evaluated. Subtracting scores from unrelated benchmarks cannot establish your process reliability. Nor can one historical result describe the failure rate of every current agent.

A demo establishes that something can work. An operational evaluation asks when it fails, how you will notice and what recovery requires.

Design a representative trial

Include ordinary cases and real exceptions: missing information, duplicate documents, unavailable tools and requests beyond the agent's scope. Use authorised data and an environment where changes can be inspected.

Repeat tasks. Include cases the agent should complete and cases where it should stop to request information.

MeasureQuestion it answers
Correct end-to-end tasksDoes it complete the whole job?
Undetected errorsWhat could proceed without review?
Human interventionsHow much supervision does it require?
Review and correction timeDoes it reduce the total work?
Actions beyond permissionDoes it respect its boundaries?

Compare with the current process using the same acceptance criteria. A faster draft may need enough review to remove the overall gain.

Match permissions to the consequences

Start with read-only access or draft changes where practical. Separate credentials and restrict tools to the operations required.

For actions that are difficult to reverse, add validation before execution. Reviewers need to see what will change and the evidence behind the decision. An approval button without that context offers little assurance.

Treat instructions inside external email, pages and documents as content to analyse, not authority to expand the task. Anthropic's work documents prompt injection as an ongoing concern for browser agents.

When to expand automation

Expand when your tests cover important exceptions and you can detect failures, stop execution and recover the previous state. Keep the test cases so you can repeat them after changing models, tools or instructions.

A bounded step with verifiable results and manageable review can be a useful first pilot.

For more on execution environments, see why the same model behaves differently with different tools. To assess a particular process, tell us its inputs, required outputs and review responsibilities.

Sources

Reviewed on 20 September 2026.

Checked on September 20, 2026

All articles