From n-gram overlap to semantic similarity and rubric-based judging — matching the metric to the task.
01
01
Understand the problem
metricsbleurougesemantic-similarityevals
02
02
Attempt it yourself
Sketch your approach before reading the solution — that's what interviews test.
Stuck? AI Nudge Available
Get a conceptual hint to guide your logic without spoiling the final implementation.
03
03
Study the solution
The solution is waiting
Give it an honest attempt first — then compare your thinking with the full walkthrough.
04
04
Read the code
The failure, concretely
reference = "Refunds take 5-7 business days"
a = "Expect your money back within about a week" # correct
b = "Refunds take 5-7 business months" # dangerously wrong
rouge(a, reference) # ~0.1 -> "bad answer" (paraphrase punished)
rouge(b, reference) # ~0.9 -> "great answer" (one wrong token ignored)
# modern equivalents:
semantic_sim(a, reference) # high — survives paraphrase
judge("factual_consistency", b) # fails — catches the wrong unit05
05
Join the discussion
Discussion (0)
Sign in to join the discussion.
No responses yet. Be the first to share what you think.