Beyond final-answer accuracy: step-level grading, tool-call correctness and trace analysis.
Skip to solutionKEEP THE
hardAI Engineering
How do you evaluate and debug agent trajectories?
874 views
01
Understand the problem
agent-evalstrajectoriestracingdebugging
02
Attempt it yourself
Sketch your approach before reading the solution — that's what interviews test.
Nudge consolestandby
Stuck? Beam a request up — the console returns a conceptual nudge that guides your logic without spoiling the implementation.
03
Study the solution
Evaluate at two levels. Outcome: did the agent achieve the goal (task success rate on a scenario suite, cost/steps as secondary metrics)? Process: were the tool calls correct (right tool, valid args), did it recover from errors, did it loop or go off-policy — graded per step by rules or an LLM judge over the full trace
Solution ready — 2 min read
Classified // press E to declassify
04
Read the code
Scenario eval with fixture replay
for (const scenario of suite) {
for (let run = 0; run < 3; run++) { // nondeterminism: sample, not once
const trace = await agent.run(scenario.task, {
tools: scenario.mode === "replay"
? fixtureTools(scenario.recording) // deterministic replay
: liveTools,
budget: { steps: 30, usd: 1 },
});
results.push({
id: scenario.id, run,
success: await scenario.check(trace), // programmatic outcome
steps: trace.steps.length,
cost: trace.costUsd,
process: gradeSteps(trace), // wrong-tool / bad-args / loop flags
});
}
}
report(aggregate(results)); // pass-rate ± variance, cost deltas vs baseline05
Join the discussion
Discussion (0)
Sign in to join the discussion.
No responses yet. Be the first to share what you think.
Transmission complete // awaiting log
KEEP THE
STREAK ALIVE.
Dossier 65 of 80 decoded in the AI Engineering track. One more won't hurt.