Перейти до основного вмісту
🔒 Режим прев'ю. Перші 15 уроків Foundations — безкоштовні; цей — Pro. Запусти 7-денний trial щоб відкрити редактор, AI-підказки і решту курсу. Потрібна картка, скасування в Dashboard у будь-який момент.
Почати 7-денний trial →
⚡
← Курси
›
AI Engineering with Python
›
Модуль 4 · Цикли та робочі процеси агента
›
Основи ДПО та РЛВЧ
quiz
56 / 105
🇺🇦
UA
▼
↗
Поділитись
⋯
Ще
+100 XP
📋
Завдання
📖
Теорія
🤖
AI Допомога
Завдання
🌐 shown in EN
📝 **Question:** What's the practical difference between RLHF and DPO? 📋 Pick the right answer. 💡 **Hint:** Re-read the theory above if unsure.
🎯 Тест
Питання
📝 **Question:** What's the practical difference between RLHF and DPO? 📋 Pick the right answer. 💡 **Hint:** Re-read the theory above if unsure.
A
DPO predates RLHF — Christiano's 2017 preference paper is the first DPO paper and OpenAI later wrapped a PPO loop around the same loss, branding the wrapper RLHF for InstructGPT
B
DPO drops the separate reward model and PPO loop — it directly trains the LLM to prefer 'chosen' over 'rejected' completions, making the pipeline simpler and cheaper while often matching RLHF quality
C
DPO needs roughly 10x more labeled preference pairs than RLHF to reach the same alignment quality, since the implicit reward signal in the contrastive loss is noisier than a fitted reward model
D
DPO and RLHF target totally different problems — DPO post-trains the embedding head while RLHF rewrites the policy logits, and modern stacks usually run both stages independently
Відповісти
💬
Зв'язатися з підтримкою