From n-gram overlap to semantic similarity and rubric-based judging — matching the metric to the task.
Skip to solutionKEEP THE
hardAI Engineering
Why are BLEU and ROUGE poor metrics for LLM outputs, and what replaced them?
766 views
01
Understand the problem
metricsbleurougesemantic-similarityevals
02
Attempt it yourself
Sketch your approach before reading the solution — that's what interviews test.
Nudge consolestandby
Stuck? Beam a request up — the console returns a conceptual nudge that guides your logic without spoiling the implementation.
03
Study the solution
BLEU and ROUGE score n-gram overlap with a reference text, which assumes one right wording — but LLM outputs are open-ended, and a perfect answer can share almost no n-grams with the reference (and vice versa: overlapping words can be factually wrong). Modern practice replaces them with embedding-based semantic similar
Solution ready — 2 min read
Classified // press E to declassify
04
Read the code
The failure, concretely
reference = "Refunds take 5-7 business days"
a = "Expect your money back within about a week" # correct
b = "Refunds take 5-7 business months" # dangerously wrong
rouge(a, reference) # ~0.1 -> "bad answer" (paraphrase punished)
rouge(b, reference) # ~0.9 -> "great answer" (one wrong token ignored)
# modern equivalents:
semantic_sim(a, reference) # high — survives paraphrase
judge("factual_consistency", b) # fails — catches the wrong unit05
Join the discussion
Discussion (0)
Sign in to join the discussion.
No responses yet. Be the first to share what you think.
Transmission complete // awaiting log
KEEP THE
STREAK ALIVE.
Dossier 68 of 80 decoded in the AI Engineering track. One more won't hurt.