Skip to main content
🔒 Preview mode. The first 15 Foundations lessons are free; this one is Pro. Start a 7-day trial to unlock the editor, AI hints and the rest of the curriculum. Card required, cancel any time in Dashboard.
Start 7-day trial →
⚡
← Courses
›
Data Science Applied
›
Module 6 · Deep Learning + MLOps
›
Evaluation: BLEU / ROUGE / LLM-as-judge
quiz
92 / 104
🇺🇸
EN
▼
↗
Share
⋯
More
+100 XP
📋
Task
📖
Theory
🤖
AI Help
Task
📝 **Question:** Best metric for open-ended summarization? 📋 Pick the right answer. 💡 **Hint:** Re-read the theory above if unsure.
🎯 Quiz
Question
📝 **Question:** Best metric for open-ended summarization? 📋 Pick the right answer. 💡 **Hint:** Re-read the theory above if unsure.
A
BLEU-4 with smoothing — the WMT shared-task standard and the only reproducible metric across research papers.
B
LLM-as-judge calibrated against human-labeled subset — captures semantic quality.
C
Exact-match accuracy against the reference — anything less than a literal-string match is treated as a model hallucination.
D
Macro F1 over keyword extraction — measures whether the summary covers the same named entities as the reference.
Submit answer
💬 Discussion
Be the first to ask a question or share a tip.
Sign in
to join the discussion. Reading is free.
Loading discussion…