Outcome vs. Process: Evaluating Multi-Step Agents

Outcome vs. Process: Evaluating Multi-Step Agents

Outcome vs. Process: Evaluating Multi-Step Agents

Judging only an agent's final answer misses most of what can go wrong. An agent plans, calls tools, and reasons across steps — and can reach a good answer by luck through a broken process that fails on the next input.

Evaluate the trajectory, not just the destination: outcome evaluation (was the result right?) and process evaluation (were the steps sound?) together.

The trajectory is what makes an agent an agent

A single model call has one output to judge. An agent has a trajectory — it plans, calls tools, observes results, reasons, and acts, often over many steps. That in-between is exactly what separates evaluating an agent from evaluating a single model call, and it's where the leverage and the failures both hide. If you only look at final answers, you're evaluating the agent as though it were a model, and missing the dimension that makes it an agent.

Outcome versus process

There are two complementary questions. Outcome evaluation asks whether the final result was correct — necessary, but blind to how it was reached. Process evaluation asks whether the steps were sound: did the agent plan sensibly, call the right tools, recover from errors, avoid needless loops? An agent that gets the right answer through a wrong process will eventually get a wrong answer, so process evaluation is what catches problems before they surface as failures.

A right answer from a wrong process is a latent bug.

What to check along the trajectory

  • Planning — did the agent break the task down sensibly, or thrash?
  • Tool selection — did it choose the right tools and call them correctly?
  • Error recovery — when a step failed, did it adapt, or spiral?
  • Efficiency — did it reach the goal in a reasonable number of steps, or loop and wander?
Want the whole evaluation loop on a few pages — define, measure, test, trust? Grab the free Agent Evaluation QuickStart.Download Free — Evaluation QuickStart

Step-level and end-to-end together

The strongest evaluation combines both levels. End-to-end checks that the whole agent accomplishes real tasks; step-level checks pinpoint where a failing agent goes wrong, so you can fix the specific step rather than guessing. End-to-end tells you that something broke; step-level tells you what. You want both, because each answers a question the other cannot.

This is why tracing matters

You can only evaluate a trajectory you can see. Capturing the full record of what the agent did — every plan, tool call, and intermediate result — is the precondition for process evaluation. Without it, a failing agent is a black box and you're left re-running a non-deterministic failure blind. Trajectory evaluation and tracing go together.

AI Agent Evaluation & Testing: The Complete Guide is the full reference — 40 pages, 15 chapters, 5 appendices, with a worked support-agent example and a 30-day adoption path.Get the Complete Guide