hardAI Engineering

How do you evaluate and debug agent trajectories?

874 views
01

Understand the problem

Beyond final-answer accuracy: step-level grading, tool-call correctness and trace analysis.

agent-evalstrajectoriestracingdebugging
02

Attempt it yourself

Sketch your approach before reading the solution — that's what interviews test.

Stuck? AI Nudge Available

Get a conceptual hint to guide your logic without spoiling the final implementation.

03

Study the solution

The solution is waiting

Give it an honest attempt first — then compare your thinking with the full walkthrough.

04

Read the code

Scenario eval with fixture replay
for (const scenario of suite) {
  for (let run = 0; run < 3; run++) {                 // nondeterminism: sample, not once
    const trace = await agent.run(scenario.task, {
      tools: scenario.mode === "replay"
        ? fixtureTools(scenario.recording)             // deterministic replay
        : liveTools,
      budget: { steps: 30, usd: 1 },
    });
    results.push({
      id: scenario.id, run,
      success: await scenario.check(trace),            // programmatic outcome
      steps: trace.steps.length,
      cost: trace.costUsd,
      process: gradeSteps(trace),                      // wrong-tool / bad-args / loop flags
    });
  }
}
report(aggregate(results));   // pass-rate ± variance, cost deltas vs baseline
05

Join the discussion

Discussion (0)

Sign in to join the discussion.

No responses yet. Be the first to share what you think.