Agent Trajectory: How to Read and Evaluate an AI Agent Run

ProductionEvaluationPublished Updated By Simon Budziak

An agent trajectory is the ordered record of one AI agent run: every message, tool call, and result. Evaluating it shows how the agent reached its answer, so you catch wrong tools, skipped approvals, and wasted calls that a final-answer score misses.

LangChain’s agent evaluation docs (read 6 October 2026) say evals measure an agent “by assessing its execution trajectory, the sequence of messages and tool calls it produces.” The raw material is a trace. In LangSmith’s terms, a trace collects every run behind one request, so a model call, a tool call and a second model call land in one record that LLM tracing stores and an evaluator reads in order.

How do you evaluate an agent trajectory?

LangChain’s open-source agentevals package offers two evaluators. create_trajectory_match_evaluator compares a run with a reference trajectory in one of four modes:

create_trajectory_llm_as_judge grades the run against a rubric instead, and the docs note it does not require a reference trajectory. That makes the LLM-as-a-judge route the practical start when nobody has written down the ideal path yet. Use the strictest match mode the task can tolerate: strict where order is the control, such as an approval step before a write, and superset where an extra lookup costs nothing.

How do we evaluate agent trajectories at Soba Labs?

We treat the recorded run as the evidence, never the agent’s account of it. Three habits come from our own agent work:

Traces are also a data export. In one agent we built, a sweep for a planted test value came back clean across logs, events, stored rows, error paths and HTTP responses, while the same value sat in every trace input: a tracing callback receives every graph node’s inputs. Our rule since: keep a field out of agent state unless a step consumes it. The wider rule is in why an agent’s summary is not a source.

What does trajectory evaluation catch that an outcome score misses?

It catches a wrong tool choice, repeated retries, a skipped approval, or a correct result reached through data the agent should not have touched. Outcome quality and execution quality are different measurements, and AI agent evals should score both whenever an agentic workflow can change external state through tool calls. Trace-based evaluation applies the same idea beyond agents. Failed production trajectories should become offline evaluation cases, so the next release is tested against the failures users actually hit.

Written with AI assistance and reviewed by Simon Budziak. The production notes come from systems Soba Labs builds and runs.

Frequently asked questions

What is the difference between an agent trajectory and a trace?

A trace stores every model call, tool call, and retrieval behind one request. The trajectory is the ordered sequence of messages and tool calls you read out of that trace and score.

Why evaluate the trajectory if the final answer is correct?

The agent may have reached the right answer through unauthorized data access, unnecessary calls, excessive cost, or a sequence that will fail on a slightly different case.

Do trajectory evals need a reference trajectory?

Only match evaluators do. An LLM-as-judge trajectory evaluator can grade a run against a rubric alone, and LangChain's docs say a reference trajectory is optional for it.

Summarize this page with

Train your team to build this