Agentic Trajectory

An agentic trajectory is the complete sequence of actions, observations, and decisions that an AI agent produces while working through a task from start to finish. It captures not just the final output but the full path the agent took to get there, including every tool call, file access, reasoning step, and intermediate result. In the context of agentic evaluation and benchmarking, a trajectory is the primary unit of analysis: it is what gets generated, stored, and assessed when measuring how well an agent performs on a given problem.

How it works

As an agent operates, it alternates between reasoning about its current state and taking actions in its environment. Each action, such as reading a file, executing a command, or calling an API, produces an observation that is fed back into the agent’s context. This loop continues until the agent reaches a stopping condition, either by submitting a final answer, hitting a step limit, or encountering a failure state. The resulting record of all steps, in order, constitutes the trajectory. Because agents are stateful and their environments can change as they act, two runs of the same agent on the same task may produce very different trajectories even when both reach a correct answer. This variability is why evaluating agentic systems requires running multiple trajectories per instance, a practice known as duplicity, and why trajectory isolation is essential to prevent one run’s file system changes from affecting another.

What it is used for

Agentic trajectories are used in evaluation, debugging, and system optimization. During evaluation, comparing trajectories across many runs reveals patterns in how agents succeed or fail, such as recurring dead ends, inefficient tool use, or systematic errors on certain problem types. In debugging, a trajectory provides a step-by-step record that engineers can inspect to understand why a run produced an incorrect or unexpected result. In optimization, trajectory data informs decisions about compute allocation, such as identifying which steps consume the most tokens or time, and serves as training signal for reinforcement learning approaches that reward efficient problem-solving paths. For infrastructure teams running large-scale agentic benchmarks, trajectories also drive architectural requirements: long trajectories demand resumability so that work is not lost to infrastructure failures, and branching trajectories from parallelism demand strict environment isolation.