Beyond final-answer accuracy: step-level grading, tool-call correctness and trace analysis.
01
01
Understand the problem
agent-evalstrajectoriestracingdebugging
02
02
Attempt it yourself
Sketch your approach before reading the solution — that's what interviews test.
Stuck? AI Nudge Available
Get a conceptual hint to guide your logic without spoiling the final implementation.
03
03
Study the solution
The solution is waiting
Give it an honest attempt first — then compare your thinking with the full walkthrough.
04
04
Read the code
Scenario eval with fixture replay
for (const scenario of suite) {
for (let run = 0; run < 3; run++) { // nondeterminism: sample, not once
const trace = await agent.run(scenario.task, {
tools: scenario.mode === "replay"
? fixtureTools(scenario.recording) // deterministic replay
: liveTools,
budget: { steps: 30, usd: 1 },
});
results.push({
id: scenario.id, run,
success: await scenario.check(trace), // programmatic outcome
steps: trace.steps.length,
cost: trace.costUsd,
process: gradeSteps(trace), // wrong-tool / bad-args / loop flags
});
}
}
report(aggregate(results)); // pass-rate ± variance, cost deltas vs baseline05
05
Join the discussion
Discussion (0)
Sign in to join the discussion.
No responses yet. Be the first to share what you think.