9 min read

AI agent evaluation before production: a practical evaluation harness

An agent that looked good in five demo conversations can still fail on the sixth real one. A small, repeatable evaluation harness turns quality from an impression into a report you can rerun on every change.

AI agent evaluation before production means running the agent against a fixed set of real, labeled cases and scoring the outcome, the path it took, and what it refused to do. A practical harness replays those cases on every prompt, model, or tool change, so regressions appear in a report, not in front of customers.

This guide describes how to build that harness without a research team: where the cases come from, how to run tools safely, how to grade results in layers, and how to make evaluation a release gate. It also covers what to skip when a simpler check is enough.

Why agents are harder to evaluate than ordinary software

A normal function returns the same output for the same input, so a unit test can assert an exact value. An agent does not. The same input can produce a different plan, a different number of tool calls, and differently worded answers from one run to the next, and several of those answers may be acceptable.

Agents also act. A wrong answer in a chatbot is a bad message; a wrong tool call in an agent can change a record or send an email. Evaluation therefore has to look at the path, not just the final text: which tools were called, with which arguments, in which order, and whether anything forbidden happened along the way.

Finally, correctness is often a judgment. Whether a drafted reply is good depends on tone, accuracy, and policy. That judgment has to be written down before it can be measured.

What teams often get wrong

  • Testing by impression. A few conversations in a chat window feel convincing and prove very little. Failures live in the long tail of inputs nobody typed during the demo.
  • Grading only the final answer. An agent can reach a correct answer through a forbidden tool call or a wasteful twelve-step path. Both matter in production.
  • Trusting a model grader blindly. Using a model to grade another model is useful, but only after you check its verdicts against human labels on a sample.
  • Happy-path cases only. A test set built from typical requests misses ambiguous inputs, missing data, hostile content, and requests the agent should decline.
  • Testing against live systems. Running evaluations against production tools creates real side effects and makes results depend on changing data.
  • Changing several things at once. Updating the prompt, the model, and a tool in one release makes it impossible to tell which change caused a regression.

A practical harness architecture

A useful harness has five parts: a case set, tool doubles, a runner, graders, and a report. Each part can start small.

The case set

Build cases from real history wherever possible: past tickets, requests, documents, or orders, with personal data removed or masked. For each case, record the input, the expected outcome, the tools that should and should not be called, and any constraints such as "must ask for the order number" or "must not issue a refund". Add categories deliberately:

  • Typical requests that should succeed.
  • Edge cases with missing, conflicting, or stale data.
  • Requests the agent should decline or escalate.
  • Adversarial inputs, such as documents or messages containing instructions aimed at the agent.
  • Cases that previously failed in production, added every time an incident happens.

A few dozen well-chosen cases per workflow are enough to start. Grow the set from real failures rather than inventing hundreds of synthetic ones up front.

Tool doubles

Replace real tools with doubles that return recorded or scripted responses and log every call. Reads return fixtures captured from real systems; writes record what the agent tried to do without doing it. This makes runs repeatable, safe, and cheap, and it lets you script failures such as timeouts or permission errors.

The runner

Run each case several times, because agents are not deterministic. A case that passes in some trials and fails in others is a real finding, not noise to average away. Record the full trace of every trial: model calls, tool calls with arguments, outputs, token usage, and timing.

Layered graders

Grade from cheapest and most reliable to most expensive:

  • Deterministic checks: output matches the required schema, only allowed tools were called, no forbidden action occurred, step count and token use stayed within limits, required fields are present.
  • Reference checks: extracted values, chosen categories, or proposed record changes match the labeled expected outcome.
  • Rubric grading by a model: for open-ended text, a grader model scores against a written rubric. Calibrate it by comparing its scores with human labels on a sample, and revisit that calibration when you change the grader model.
  • Human review: a person reviews a sample of results, especially failures and borderline scores, and turns new failure patterns into new cases.

The report

Report results per case and per category, compared with the previous run. The most useful view is the list of cases that changed status, with the traces side by side. Aggregate scores are for trends; individual diffs are for decisions.

Implementation considerations

  • Version everything together. Store the prompt, model identifier, tool definitions, and case set version with every evaluation run, so results are reproducible.
  • Make it a gate. Run the harness in continuous integration for prompt, model, and tool changes, and block release when critical cases fail. Treat a prompt change like a code change.
  • Budget the evaluation. Repeated trials across many cases cost money and time. Run a fast subset on every change and the full set before release.
  • Protect the data. Historical cases often contain personal or commercial information. Mask it, restrict access, and keep the case set out of public repositories.
  • Keep the harness independent. A harness that does not depend on your agent framework lets you compare frameworks or models on identical cases.
  • Monitor after launch. Production traces, reviewer edits, and user complaints are the best source of new cases. Evaluation does not stop at release.

Trade-offs

Deterministic checks are cheap and trustworthy but only cover what you can specify precisely. Model graders cover open-ended quality but add cost and their own error. Human review is the most accurate and the least scalable. A good harness uses all three in proportion.

More trials per case give a clearer picture of reliability and cost more. More cases give broader coverage and take longer to label. Start narrow and deep on the workflows with the highest cost of failure.

Tool doubles make runs safe and repeatable but can drift from reality. Refresh fixtures from real systems periodically, and keep a small set of end-to-end checks against a sandbox environment.

Lessons from ImadDhin work

These are code-level observations from this portal's repository, not client outcomes.

  • Structure first, behavior next. A common finding in inspected agent codebases is that automated tests cover the surrounding scripts and integrations, such as attribution handling or CRM field mapping, while agent behavior relies on structural guards alone. Guards are necessary but not sufficient; a behavioral evaluation suite is what turns prompt rules into verified behavior, and building one is what this article describes.
  • Structural guards are directly testable. Tool inputs are validated with typed schemas, CRM writes are limited to allowlisted fields, runs are capped at a fixed number of graph steps, and gated writes cannot proceed without an approved record. Each of these can be covered by deterministic tests without calling a model at all.
  • The run log is a ready trace format. Every run writes a sequenced event log of steps, tool starts and completions, approvals, and model usage. That is exactly the trace a harness needs to replay and grade runs.
  • Prompt rules are hypotheses. System instructions tell the models never to invent tool results, prices, or approvals, and to treat documents as data. Those are claims to test with adversarial cases, not guarantees.
  • Error paths deserve cases too. Chat falls back across several model providers and returns a visitor-safe message when all fail. Scripted provider failures belong in the case set so the fallback behavior is verified, not assumed.

Common mistakes to test for

  • A document containing instructions to call a tool or reveal system rules.
  • A request that should be declined, such as a refund the agent has no authority to issue.
  • Missing required information, where the right behavior is to ask rather than guess.
  • A tool that times out or returns an error, where the right behavior is to stop or escalate.
  • Two similar inputs that should lead to different actions.
  • A case that previously caused an incident, kept permanently in the set.

When a simpler solution is better

If the model does a single classification or extraction step inside a fixed workflow, you do not need a full harness. A spreadsheet of labeled examples and a short script that computes agreement is enough. If volume is tiny and every output is reviewed by a person anyway, that review is your evaluation for now; record the outcomes so they become test cases later. Build the full harness when the agent takes actions, runs at volume, or changes often.

Evaluation is also what makes the other controls trustworthy. Approval gates, described in human-in-the-loop agents, generate labeled data every day, and the security testing in securing agents belongs in the same regression suite.

Make quality a report, not an impression

Start with a few dozen real cases, tool doubles, and deterministic checks, then add model grading and human review where they earn their cost. If you want an evaluation harness for an agent you are building or already running, see AI agent development or bring a sample of real cases to a 30-minute call.

Frequently asked questions

How many test cases does an AI agent need?

Start with a few dozen well-chosen real cases per workflow, covering typical, edge, decline, and adversarial categories. Grow the set from production failures and reviewer edits rather than inventing large synthetic sets up front.

Can I use an LLM to grade my agent?

Yes, for open-ended outputs, with a written rubric. Check the grader's verdicts against human labels on a sample first, and recheck whenever you change the grader model.

Should evaluations run against real systems?

Mostly no. Use tool doubles with recorded fixtures so runs are safe, repeatable, and cheap, and keep a small number of end-to-end checks against a sandbox environment.

How often should I rerun agent evaluations?

On every change to the prompt, model, or tools, with a fast subset in continuous integration and the full set before each release. Add new cases whenever production reveals a failure.

What should an agent evaluation measure besides the final answer?

The path: which tools were called with which arguments, whether any forbidden action occurred, how many steps and tokens were used, and whether the agent asked, declined, or escalated when it should have.

Evaluate your agent before your customers do

Bring a sample of real cases; leave with a plan for your first harness.

Book a 30-minute call

Evaluation harnesses, adversarial testing, and regression gates for agents.

See AI agent development

Keep reading