hardAI Engineering

Why are BLEU and ROUGE poor metrics for LLM outputs, and what replaced them?

765 views
01

Understand the problem

From n-gram overlap to semantic similarity and rubric-based judging — matching the metric to the task.

metricsbleurougesemantic-similarityevals
02

Attempt it yourself

Sketch your approach before reading the solution — that's what interviews test.

Stuck? AI Nudge Available

Get a conceptual hint to guide your logic without spoiling the final implementation.

03

Study the solution

The solution is waiting

Give it an honest attempt first — then compare your thinking with the full walkthrough.

04

Read the code

The failure, concretely
reference = "Refunds take 5-7 business days"

a = "Expect your money back within about a week"   # correct
b = "Refunds take 5-7 business months"             # dangerously wrong

rouge(a, reference)   # ~0.1  -> "bad answer"      (paraphrase punished)
rouge(b, reference)   # ~0.9  -> "great answer"    (one wrong token ignored)

# modern equivalents:
semantic_sim(a, reference)          # high — survives paraphrase
judge("factual_consistency", b)     # fails — catches the wrong unit
05

Join the discussion

Discussion (0)

Sign in to join the discussion.

No responses yet. Be the first to share what you think.