LangChain’s agent evaluation docs (read 6 October 2026) say evals measure an agent “by assessing its execution trajectory, the sequence of messages and tool calls it produces.” The raw material is a trace. In LangSmith’s terms, a trace collects every run behind one request, so a model call, a tool call and a second model call land in one record that LLM tracing stores and an evaluator reads in order.
How do you evaluate an agent trajectory?
LangChain’s open-source agentevals package offers two evaluators. create_trajectory_match_evaluator compares a run with a reference trajectory in one of four modes:
strict: the same tool calls in the same orderunordered: the same tool calls in any ordersubset: no tools beyond the referencesuperset: at least the reference tools, extras allowed
create_trajectory_llm_as_judge grades the run against a rubric instead, and the docs note it does not require a reference trajectory. That makes the LLM-as-a-judge route the practical start when nobody has written down the ideal path yet. Use the strictest match mode the task can tolerate: strict where order is the control, such as an approval step before a write, and superset where an extra lookup costs nothing.
How do we evaluate agent trajectories at Soba Labs?
We treat the recorded run as the evidence, never the agent’s account of it. Three habits come from our own agent work:
- To prove a tool ran, we match the call entry and its result payload, not the tool’s name. In one coding-agent session log the tool’s name appeared 32 times and the tool was called zero times, because the prompt named it too.
- We read the model and reasoning effort from the session log, not from the flags we passed. A reasoning-effort setting written in the wrong syntax is silently ignored, and the run lands at medium effort with no warning.
- When a long run dies at its final summary step, the log still holds every finding produced on the way, so we recover from the trajectory instead of relaunching.
Traces are also a data export. In one agent we built, a sweep for a planted test value came back clean across logs, events, stored rows, error paths and HTTP responses, while the same value sat in every trace input: a tracing callback receives every graph node’s inputs. Our rule since: keep a field out of agent state unless a step consumes it. The wider rule is in why an agent’s summary is not a source.
What does trajectory evaluation catch that an outcome score misses?
It catches a wrong tool choice, repeated retries, a skipped approval, or a correct result reached through data the agent should not have touched. Outcome quality and execution quality are different measurements, and AI agent evals should score both whenever an agentic workflow can change external state through tool calls. Trace-based evaluation applies the same idea beyond agents. Failed production trajectories should become offline evaluation cases, so the next release is tested against the failures users actually hit.