Skip to solution
hardAI Engineering

How do you evaluate and debug agent trajectories?

874 views
01

Understand the problem

Beyond final-answer accuracy: step-level grading, tool-call correctness and trace analysis.

agent-evalstrajectoriestracingdebugging
02

Attempt it yourself

Sketch your approach before reading the solution — that's what interviews test.

Nudge consolestandby

Stuck? Beam a request up — the console returns a conceptual nudge that guides your logic without spoiling the implementation.

03

Study the solution

Evaluate at two levels. Outcome: did the agent achieve the goal (task success rate on a scenario suite, cost/steps as secondary metrics)? Process: were the tool calls correct (right tool, valid args), did it recover from errors, did it loop or go off-policy — graded per step by rules or an LLM judge over the full trace

Solution ready — 2 min read

Classified // press E to declassify

04

Read the code

Scenario eval with fixture replay
for (const scenario of suite) {
  for (let run = 0; run < 3; run++) {                 // nondeterminism: sample, not once
    const trace = await agent.run(scenario.task, {
      tools: scenario.mode === "replay"
        ? fixtureTools(scenario.recording)             // deterministic replay
        : liveTools,
      budget: { steps: 30, usd: 1 },
    });
    results.push({
      id: scenario.id, run,
      success: await scenario.check(trace),            // programmatic outcome
      steps: trace.steps.length,
      cost: trace.costUsd,
      process: gradeSteps(trace),                      // wrong-tool / bad-args / loop flags
    });
  }
}
report(aggregate(results));   // pass-rate ± variance, cost deltas vs baseline
05

Join the discussion

Discussion (0)

Sign in to join the discussion.

No responses yet. Be the first to share what you think.

Transmission complete // awaiting log

KEEP THE
STREAK ALIVE.

Dossier 65 of 80 decoded in the AI Engineering track. One more won't hurt.

Back to track