Skip to solution
hardAI Engineering

Why are BLEU and ROUGE poor metrics for LLM outputs, and what replaced them?

766 views
01

Understand the problem

From n-gram overlap to semantic similarity and rubric-based judging — matching the metric to the task.

metricsbleurougesemantic-similarityevals
02

Attempt it yourself

Sketch your approach before reading the solution — that's what interviews test.

Nudge consolestandby

Stuck? Beam a request up — the console returns a conceptual nudge that guides your logic without spoiling the implementation.

03

Study the solution

BLEU and ROUGE score n-gram overlap with a reference text, which assumes one right wording — but LLM outputs are open-ended, and a perfect answer can share almost no n-grams with the reference (and vice versa: overlapping words can be factually wrong). Modern practice replaces them with embedding-based semantic similar

Solution ready — 2 min read

Classified // press E to declassify

04

Read the code

The failure, concretely
reference = "Refunds take 5-7 business days"

a = "Expect your money back within about a week"   # correct
b = "Refunds take 5-7 business months"             # dangerously wrong

rouge(a, reference)   # ~0.1  -> "bad answer"      (paraphrase punished)
rouge(b, reference)   # ~0.9  -> "great answer"    (one wrong token ignored)

# modern equivalents:
semantic_sim(a, reference)          # high — survives paraphrase
judge("factual_consistency", b)     # fails — catches the wrong unit
05

Join the discussion

Discussion (0)

Sign in to join the discussion.

No responses yet. Be the first to share what you think.

Transmission complete // awaiting log

KEEP THE
STREAK ALIVE.

Dossier 68 of 80 decoded in the AI Engineering track. One more won't hurt.

Back to track