An AI agent can produce a convincing final message while leaving the underlying task incomplete or changing the wrong record. That is why production evaluation must inspect behavior and outcomes, not just the quality of the prose. A good evaluation answers a practical question: can this configuration perform this class of work within the boundaries the business has authorized?
Consider an illustrative support agent that investigates a billing discrepancy and prepares a correction for review. It needs to identify the right account, gather relevant evidence, apply the current policy, and stop before an unapproved change. The final explanation matters, but so do the records accessed, tool calls made, and proposal stored in the application.
This guide lays out a repeatable evaluation process for teams moving beyond demos. It covers task selection, grading, repeated trials, security cases, release decisions, and production feedback. The goal is evidence that supports a bounded deployment decision, not a single score used to claim an agent is universally reliable.
Key takeaways
- Define success through observable task outcomes, allowed actions, and required evidence.
- Use representative failures and ambiguous cases, not only clean demonstration prompts.
- Combine deterministic checks, calibrated review, and repeated trials where behavior varies.
- Make evaluation results a release input and keep production incidents feeding the test set.
Turn the product promise into testable tasks
Start with the work the agent is supposed to perform. Break a broad promise such as resolve billing questions into task classes: explain a charge, find a duplicate invoice, draft a correction, or escalate an account mismatch. Each class needs its own completion conditions and permissions.
Write a task specification containing the initial state, user request, available tools, allowed scope, expected outcome, and unacceptable actions. For the billing example, the agent may read records for one organization and prepare a proposal but may not issue a refund without approval.
Include what the agent should do when it cannot complete the task. Asking for a missing identifier or escalating conflicting evidence can be correct behavior. If the evaluation rewards only completed answers, it can encourage the system to guess rather than acknowledge a genuine boundary.
Keep task specifications understandable to product owners and reviewers. They should describe business outcomes and evidence, not only model-specific prompt details. This makes the evaluation useful even when the underlying model or orchestration framework changes.
Build a dataset that represents actual work
Collect examples from observed workflows, support patterns, and product research, using synthetic or appropriately prepared data to avoid unnecessary exposure. Include ordinary cases, difficult cases, and cases that should be rejected. Do not let the dataset become a collection of prompts chosen because the agent already handles them well.
Segment by factors that may change performance: task type, data volume, ambiguity, user role, organization context, and dependency availability. An overall success rate can hide a weak category that matters disproportionately to the business.
Reserve some examples for comparison outside routine prompt development. Repeatedly tuning against the same visible cases can produce an agent that performs well on familiar tests without improving general behavior. Keep a stable benchmark and add newly discovered failures separately.
Document where examples came from and why they matter. A small, carefully chosen set with clear expected outcomes is more useful than a large dataset of loosely labeled conversations. Quality of coverage should guide growth, especially early in the evaluation program.
Inspect state and tool behavior, not only the final answer
Anthropic’s agent-evaluation guidance emphasizes tasks, trials, graders, and the distinction between a transcript and the resulting environment state. That distinction is central to practical testing: an agent saying it created a correct draft is not proof that the correct draft exists.
After a run, inspect the relevant application records. Check the target account, field values, number of created objects, and any unintended changes. For read-only tasks, verify that the agent accessed only permitted resources and used evidence relevant to the question.
Capture a structured trace of tool calls and results with sensitive data handled appropriately. The trace should help explain failures without requiring reviewers to read an enormous raw conversation. Highlight denied actions, retries, budget limits, and state transitions.
Define forbidden outcomes explicitly. A task that produces a helpful answer while exposing another organization’s data must fail. Safety and authorization properties should not disappear inside an average score where good writing compensates for unacceptable behavior.
Use deterministic checks wherever the outcome is concrete
Many important requirements can be checked without another model. Confirm that a record belongs to the correct organization, an amount matches the approved proposal, a required field is present, or no write tool was called during a read-only task.
Write these checks against the behavior you intend, not the agent’s implementation details. A test should not require one exact sequence of harmless reads if several sequences can produce the correct result. It should enforce the outcome and boundaries that matter.
Be careful with overly rigid text matching. A correct explanation can be phrased in many ways. Reserve exact checks for identifiers, required facts, structured fields, and other elements where equality is meaningful. Use a different review method for clarity or completeness.
Version the check logic and fixtures. If a policy changes, update the expected result deliberately and preserve the reason. Otherwise a changing test harness can make results appear better or worse without any corresponding change in the agent.
Calibrate human and model-based grading
Some qualities require judgment: whether an explanation is clear, whether uncertainty is communicated appropriately, or whether the response addresses the user’s real goal. Create a rubric with examples of acceptable, borderline, and unacceptable outputs.
Have reviewers grade a shared sample and discuss disagreements. This exposes ambiguous criteria before they are used to judge many runs. The aim is not perfect agreement on every stylistic detail but a consistent understanding of the decisions that affect users.
A model-based grader can help prioritize or scale review, but it should be checked against human-reviewed examples. Inspect cases where it disagrees, and avoid giving it access to irrelevant signals that could bias the score. A polished answer should not automatically receive a high correctness rating.
Keep separate dimensions visible. Evidence quality, task completion, permission compliance, clarity, and cost answer different questions. Combining them into one number may be convenient for a dashboard, but release decisions should still examine the dimensions whose failure has serious consequences.
Run repeated trials and analyze failure patterns
Agent behavior can vary across runs even with the same task. Use repeated trials for representative cases, particularly where the workflow involves several decisions or tool calls. Record the configuration so the team can distinguish natural variation from an actual version change.
Report the distribution of outcomes and the sample size. A single successful run is weak evidence for a task that sometimes fails in consequential ways. Avoid precise reliability claims based on a small or unrepresentative set.
Classify failures by cause: wrong task interpretation, missing evidence, retrieval failure, tool misuse, policy misunderstanding, state-management error, or poor recovery. This classification helps identify the appropriate fix. More prompt instructions will not repair a connector that returns the wrong account.
Review successful runs too. They may reveal unnecessary tool calls, excessive latency, or an accidental dependency on information the agent should not have used. A result can be correct while the process remains too expensive, fragile, or difficult to audit for production use.
Test permissions, hostile content, and changing conditions
Include users with different roles, revoked access, and records from other organizations. Test whether the application enforces boundaries even when the agent proposes an unauthorized action. The model should never be the sole access-control mechanism.
Place conflicting instructions in documents, messages, and tool results. The agent should treat retrieved material as data rather than authority to change its task or reveal secrets. Evaluate actual tool activity and data exposure, not only whether the final answer repeats the malicious instruction.
Test conditions that change during a task. An approval can expire, a record can be updated, and an account can lose access. The workflow should revalidate relevant conditions before a consequential action instead of assuming the initial state remains true indefinitely.
Include resource-abuse scenarios such as repeated searches, large inputs, and loops that fail to converge. Verify that step, time, concurrency, and spending limits are enforced by the application and produce an understandable incomplete result.
Exercise recovery and operational controls
Simulate a tool timeout after an action may have succeeded. Check whether the agent and workflow can discover the resulting state before retrying. Duplicate tickets, messages, or commercial actions can arise when a timeout is treated as proof that nothing happened.
Test partial completion and cancellation. If the task stops after creating a draft but before attaching evidence, the result should describe that state and support recovery. Cancellation should prevent further steps without falsely claiming to undo effects already completed.
Verify capability-level shutdown. Operators should be able to disable a failing write tool or a problematic task class without necessarily removing every useful function. The fallback should be documented and visible to the people handling affected work.
Run an incident rehearsal with the actual operators. Ask them to locate a failed task, identify the configuration, inspect the relevant evidence, pause the capability, and resume the ordinary workflow. Operational readiness is part of evaluation because production failures require more than a model score to resolve.
Define release gates according to consequence
Set release criteria before reviewing the final results. Identify which failures are unacceptable, which can be handled through a limited rollout, and which require a manual step. Different task classes may have different gates because their consequences differ.
A read-only summarizer and an agent that changes customer accounts should not share one generic acceptance threshold. For consequential actions, permission failures, incorrect targets, and duplicate effects may block release even if most other tasks succeed.
Compare the proposed version with the current baseline. Look for regressions by category as well as overall improvement. A new model might improve complex reasoning while making simple tool selection less reliable or increasing cost substantially.
Document the release decision, known limitations, rollout group, monitoring, and rollback conditions. This makes the decision accountable and gives operators a clear reference. An evaluation should support a concrete deployment scope rather than end with an unexplained green badge.
Turn production feedback into a durable evaluation loop
Collect user feedback and operational failures with appropriate privacy controls. Link reports to the task type, configuration version, and relevant system state. A vague complaint that the agent was wrong is difficult to act on without context.
Convert recurring or consequential failures into reproducible cases. Remove unnecessary sensitive information while preserving the relationships that caused the problem. Add the case to the appropriate regression set and verify the fix against related tasks.
Review the evaluation program itself as the product evolves. New tools, customer groups, and policies create new failure modes. Retire obsolete cases deliberately, retain a stable comparison core, and keep business owners involved in defining successful outcomes. Reliable agents require an ongoing feedback system around them; evaluation is the mechanism that makes improvements and limitations visible.
Frequently asked questions
How many evaluation cases do we need?
There is no universal number. Begin with enough cases to cover the important task types, permissions, failure modes, and ambiguous situations. Expand based on production evidence and the consequences of errors. A large set of repetitive easy cases can provide less assurance than a smaller set with meaningful coverage.
Can another AI model grade everything?
No. Use deterministic checks for concrete state and permission requirements, and calibrated human review for important judgment. Model-based grading can assist, but it needs validation against reviewed examples and ongoing inspection. It should not be the only evidence for consequential business actions.
What should block a release?
Failures that violate the agreed authorization, data, or business boundaries should receive particular attention. Define blocking conditions for each task class before testing. A high average score should not compensate for accessing the wrong customer record or executing an unapproved action.
Do we need to rerun evaluations after changing a prompt?
Yes, at a scope appropriate to the change. Prompts, models, tools, retrieval, and source policies all influence behavior. Run affected cases and the stable regression set, compare category-level results, and inspect unexpected changes. Small textual edits can have effects beyond the example that motivated them.
Working through this in your product?
We can help you turn these decisions into a practical plan and working software.
Production AI agent engineeringTalk to the team
