ThinkDeck

How to Evaluate AI Agents: Test the Path, Not Just the Answer

AI agent evaluation beyond final answers: trajectory checks, tool-use metrics, eval sets from real tasks, and staged rollout from sandbox to production.

AI AgentsBy Published Updated 4 min read

Testing a normal feature is straightforward: given this input, expect that output. Testing an AI agent is harder, because two runs can reach the same answer by completely different routes, and one of those routes might have emailed the wrong customer along the way.

That's why AI agent evaluation has to look at the path as well as the destination. This guide covers what to measure, how to build an evaluation set, and how we roll agents out at ThinkDeck so problems surface before your customers find them.

Why checking the final answer isn't enough

Imagine an agent asked to cancel a subscription. Run A looks up the account, confirms the plan, cancels it, and sends a confirmation. Run B cancels the wrong plan, notices the error, reactivates it, cancels the right one, and sends a confirmation. Both end with the correct result. Only one of them should be allowed near production.

The full sequence of reasoning steps and tool calls an agent takes is called its trajectory. Evaluating trajectories is what separates agent testing from ordinary output checks.

What to look for in a trajectory

  • Tool choice. Did it pick the right tool for each step, or reach for something loosely related?
  • Tool arguments. Right customer ID, right amount, right date range?
  • Efficiency. Did it take five steps where two would do? Extra steps cost money and add risk.
  • Recovery. When a tool failed or returned something unexpected, did it adapt sensibly or plough on?
  • Clarifying questions. Did it ask when the request was genuinely ambiguous, and not ask when the answer was already available?
  • Policy. Did it stay inside its rules: approval before refunds, no contact with excluded customers, no actions beyond its permissions?

Four layers of agent testing

1. Unit tests for tools

Every tool is ordinary code and should be tested like it. If issue_refund mishandles currency, no amount of prompt tuning will fix the agent.

2. Scenario tests

Scripted tasks with known correct outcomes, run end to end against test data or sandboxed systems. Include the awkward ones: missing information, conflicting records, a tool that times out, a user who changes their mind.

3. Trajectory review

Automated checks on the path: required steps happened, forbidden steps didn't, step count stayed under budget. A model can grade softer qualities (was the clarifying question reasonable?) against a written rubric, but spot-check those grades yourself. Model judges have blind spots of their own.

4. Human review of a sample

Someone who knows the business reads a sample of real runs every week. This is where you catch the failures nobody thought to write a test for.

Metrics worth tracking

MetricWhat it tells you
Task success rateHow often the agent achieves the goal correctly
Policy violation rateHow often it breaks a rule, even if the result looked fine
Escalation rateHow often it hands off to a person (too high is useless, too low can be risky)
Steps per taskEfficiency, and an early warning for loops
Cost per taskModel and tool spend for one completed task
LatencyHow long users or downstream systems wait

Cost and latency per step come straight from AiKey, the AI gateway we route agent model calls through, so they sit alongside quality metrics instead of in a separate billing report.

Building the evaluation set

  1. Start from real work. Pull 30–100 historical examples of the task and record the correct outcome for each.
  2. Add the edge cases your team already knows about: the customer with two accounts, the invoice in the wrong currency.
  3. Grow it from production. Every failure found in review becomes a new test case, so the same mistake can't quietly return.
  4. Re-run on every change. New prompt, new tool, new model version: run the full set before shipping.

Staged rollout: sandbox, shadow, canary, production

StageWhat happensMove on when
SandboxAgent runs against test data onlyEval set passes at your target rate
ShadowAgent runs on real tasks; a person approves every actionApprovals are consistently a formality
CanaryAgent acts on its own for a small share of tasksMetrics hold steady against the human baseline
ProductionFull rollout, with monitoring and sampling(Keep reviewing; this stage never really ends)

Each stage tests something different. Sandbox proves the logic. Shadow proves it on messy real inputs. Canary proves it holds up without a safety net. More on the infrastructure side in deploying AI agents to production.

Evaluation is part of the build

In our AI agent development projects, the evaluation set is written before the agent is, and the dashboards ship with it. If you're starting out, read how to build an AI agent for the build steps, and context engineering for fixing the failures your evals uncover.

// work_with_thinkdeck

AI agent development services

We scope, build, and monitor production AI agents for startups, with guardrails and evaluation built in.

Explore AI agent development services

Frequently asked questions

What is trajectory evaluation for AI agents?

+

Checking the full sequence of reasoning steps and tool calls an agent took, not just its final answer: tool choice, arguments, efficiency, error recovery, and whether it followed its rules.

How many test cases does an AI agent evaluation set need?

+

Start with 30–100 real examples with known correct outcomes, including known edge cases, then add every failure found in production so the set grows with the agent.

Can an LLM grade another AI agent's work?

+

Yes, for qualities that are hard to check with code, if you give it a clear rubric. Spot-check its grades regularly, because model judges can miss the same kinds of errors the agent makes.

What is a canary rollout for an AI agent?

+

Letting the agent act on its own for a small share of real tasks while the rest stay with the existing process, then comparing results before expanding.

// next_step

Tell us what you're building.
We'll tell you how fast we can ship it.

contact@thinkdeck.site