Прескочи на главни садржај
🔒 Режим прегледа. Првих петнаест Foundations лекција је бесплатно; ова је Pro. Покрените 7-дневни trial да откључате едитор, AI савете и остатак курса. Картица је обавезна, можете отказати у било ком тренутку у Dashboard.
Покрени 7-дневни trial →
⚡
← Kursevi
›
AI Engineering with Python
›
Модул 4 · Агент Лоопс & Воркфловс
›
Основе ДПО и РЛХФ
quiz
56 / 105
🇷🇸
SR
▼
↗
Подели
⋯
Више
+100 XP
📋
Zadatak
📖
Teorija
🤖
AI Pomoć
Zadatak
🌐 shown in EN
📝 **Question:** What's the practical difference between RLHF and DPO? 📋 Pick the right answer. 💡 **Hint:** Re-read the theory above if unsure.
🎯 Kviz
Pitanje
📝 **Question:** What's the practical difference between RLHF and DPO? 📋 Pick the right answer. 💡 **Hint:** Re-read the theory above if unsure.
A
DPO predates RLHF — Christiano's 2017 preference paper is the first DPO paper and OpenAI later wrapped a PPO loop around the same loss, branding the wrapper RLHF for InstructGPT
B
DPO drops the separate reward model and PPO loop — it directly trains the LLM to prefer 'chosen' over 'rejected' completions, making the pipeline simpler and cheaper while often matching RLHF quality
C
DPO needs roughly 10x more labeled preference pairs than RLHF to reach the same alignment quality, since the implicit reward signal in the contrastive loss is noisier than a fitted reward model
D
DPO and RLHF target totally different problems — DPO post-trains the embedding head while RLHF rewrites the policy logits, and modern stacks usually run both stages independently
Odgovori
💬
Kontaktiraj podršku