Deterministic checks vs judged qualities — and why good eval suites layer both.
Skip to solutionKEEP THE
mediumAI Engineering
When do you use code-based graders versus model-based graders?
969 views
01
Understand the problem
gradersevalscode-gradersllm-as-judge
02
Attempt it yourself
Sketch your approach before reading the solution — that's what interviews test.
Nudge consolestandby
Stuck? Beam a request up — the console returns a conceptual nudge that guides your logic without spoiling the implementation.
03
Study the solution
Use code graders for anything objectively checkable: exact/contains matches, schema validity, numeric tolerance, code that compiles and passes tests, latency/cost budgets — they are deterministic, free and unarguable. Use model graders for qualities code cannot verify: relevance, faithfulness to context, tone, complete
Solution ready — 2 min read
Classified // press E to declassify
04
Read the code
Layered grading, reported separately
async function grade(c: EvalCase, out: Output) {
const gates = {
schema: Schema.safeParse(out.json).success,
citations: citationsResolve(out.text, c.contextIds),
banned: !BANNED.some((w) => out.text.includes(w)),
latency: out.ms < 3000,
};
if (Object.values(gates).some((v) => !v)) return { gates, judged: null }; // fail fast
const judged = {
faithfulness: await judge.score("faithfulness", c, out), // survivors only
tone: await judge.score("tone", c, out),
};
return { gates, judged };
}05
Join the discussion
Discussion (0)
Sign in to join the discussion.
No responses yet. Be the first to share what you think.
Transmission complete // awaiting log
KEEP THE
STREAK ALIVE.
Dossier 23 of 80 decoded in the AI Engineering track. One more won't hurt.