Skip to main content

AI Quality Engineering

How to evaluate AI agents before release

AI agents need more than output checks: task completion, tool-use accuracy, escalation behaviour, safety, cost, and auditability all need evaluation before production.

CAPTIVOLT INSIGHTS · Published · Updated · 5 min read

Executive summary

AI agents need more than output checks. Enterprises must evaluate task completion, tool-use accuracy, instruction adherence, escalation behaviour, safety, cost, latency, and auditability before agents enter production. An agent that writes beautiful prose but calls the wrong API is not a quality problem. It is an incident.

The problem

Traditional QA asks: is the output correct? Agents demand a harder set of questions: did it complete the task, did it choose the right tool with the right inputs, did it follow instructions when instructions conflicted with convenience, did it escalate when it should have, and can you prove all of this afterwards? Most teams discover these questions after the first production incident, which is the most expensive possible time to discover them.

The Captivolt thesis

An agent is ready for release when its behaviour has been measured against thresholds written before it was built, not when a demonstration went well.

A practical framework

  1. 01

    Task completion: did the agent achieve the intended outcome across representative scenarios, including edge cases?

  2. 02

    Tool-use accuracy: right tool, right parameters, right sequence, measured, not assumed.

  3. 03

    Instruction adherence: does the agent respect constraints under pressure, ambiguity, and adversarial input?

  4. 04

    Escalation behaviour: does it hand off to humans at the defined thresholds, every time?

  5. 05

    Safety and policy checks: prompt injection resistance, unsafe-output handling, data-boundary respect.

  6. 06

    Cost and latency: per-task economics measured against thresholds before scale, not after the invoice.

  7. 07

    Auditability: every decision, tool call, and escalation logged in a form governance can use.

Going further

Before any evaluation runs

The questions this article asks are only answerable if these were decided before the build. Most teams decide them after the first incident.

Name the in-scope tasks

Which tasks the agent is for, and which it should decline. An agent evaluated against everything has been evaluated against nothing.

Write the escalation thresholds

The conditions under which it must hand off, stated as rules a test can check, before anybody builds the behaviour.

Agree the cost ceiling

Cost per task, agreed as a limit before scale. After the invoice it is a negotiation, not an evaluation.

Name what must be auditable

Which decisions, tool calls and escalations governance will ask for, so the logging is designed rather than retrofitted.

Name the accountable owner

Somebody who can hold a release. Without that authority the evaluation produces a report, not a decision.

Building the evaluation set

The dataset decides what the evaluation can see, and the failures that matter are the ones a happy-path set never contains.

Representative tasks

Real inputs from the workflow, including the badly phrased and incomplete ones.

Tool-use traps

Cases where the obvious tool is wrong, or a required input is missing and has to be asked for rather than guessed.

Conflicting instructions

The user asks for something the policy forbids. Adherence is only tested when compliance is inconvenient.

Escalation triggers

Cases designed to cross each threshold, so escalation is tested as behaviour rather than asserted as design.

Adversarial input

Instructions injected into retrieved content and tool responses, not only into the prompt.

Examples

Worked through elsewhere on this site.

Reference flows and published architectures rather than client runs, each described as what it is.

Checklist

Agent release checklist

Every item is a yes-or-no question somebody can answer before an agent enters production. The thresholds are yours to set; the checklist asks whether you set them.

Before evaluation

  • In-scope tasks are written down, including what the agent should decline
  • Escalation thresholds are stated as rules a test can check
  • A cost-per-task ceiling is agreed
  • The decisions, tool calls and escalations that must be auditable are named
  • An accountable owner can hold the release

Task completion

  • The agent completes representative tasks, including edge cases
  • Incomplete or ambiguous inputs are handled by asking, not by guessing
  • Side effects match the task, and none occur that should not

Tool use

  • The right tool is chosen, including when the obvious one is wrong
  • Tool inputs validate against their schema
  • Tools are called in the right sequence and no more often than the work requires

Instruction adherence

  • Constraints hold when they conflict with what the user asked for
  • Constraints hold under ambiguous and adversarial input

Escalation

  • Every defined threshold triggers a handoff in testing
  • Handoffs carry enough context for a person to act on them

Safety and policy

  • Instructions injected into retrieved content and tool responses are ignored
  • Unsafe requests are refused, and permitted requests are not over-refused
  • Data boundaries hold for users with different entitlements

Cost and latency

  • Cost per task is measured and within the agreed ceiling
  • Latency is measured under realistic load, not on a quiet afternoon

Auditability

  • Every decision, tool call and escalation is logged in a form governance can use
  • A past run can be reconstructed from the record alone

A Markdown file made in your browser. No email address, and nothing to sign up for.

Practical implications

  • Agent evaluation is multi-dimensional; output checks alone are negligent.
  • Define escalation thresholds before build, then test against them.
  • Cost per task is an evaluation dimension, not an afterthought.
  • If you cannot audit it, you cannot deploy it in a regulated environment.

What leaders should do

  1. Name the accountable owner who can hold a release before the build starts.
  2. Require escalation thresholds and a cost ceiling to be written as rules a test can check.
  3. Ask to see the evaluation set: it should hold badly phrased, adversarial and out-of-scope cases, not only the happy path.
  4. Treat a passing evaluation as dated evidence, and have it re-run on every change to model, prompt, tools or data.

About this article

Author
Captivolt Insights
Published
· updated