Ugrás a fő tartalomra
🔒 Előnézet mód. Az első tizenöt Foundations lecke ingyenes; ez Pro. Indíts 7 napos trial-t, hogy feloldd a szerkesztőt, az AI tippeket és a tananyag többi részét. Kártya szükséges, bármikor lemondhatod a Dashboard-ban.
7 napos trial indítása →
⚡
← Kurzusok
›
AI Engineering with Python
›
4. modul · Ügynökhurkok és munkafolyamatok
›
DPO és RLHF alapok
quiz
56 / 105
🇭🇺
HU
▼
↗
Megosztás
⋯
Több
+100 XP
📋
Feladat
📖
Elmélet
🤖
AI Segítség
Feladat
🌐 shown in EN
📝 **Question:** What's the practical difference between RLHF and DPO? 📋 Pick the right answer. 💡 **Hint:** Re-read the theory above if unsure.
🎯 Kvíz
Kérdés
📝 **Question:** What's the practical difference between RLHF and DPO? 📋 Pick the right answer. 💡 **Hint:** Re-read the theory above if unsure.
A
DPO predates RLHF — Christiano's 2017 preference paper is the first DPO paper and OpenAI later wrapped a PPO loop around the same loss, branding the wrapper RLHF for InstructGPT
B
DPO drops the separate reward model and PPO loop — it directly trains the LLM to prefer 'chosen' over 'rejected' completions, making the pipeline simpler and cheaper while often matching RLHF quality
C
DPO needs roughly 10x more labeled preference pairs than RLHF to reach the same alignment quality, since the implicit reward signal in the contrastive loss is noisier than a fitted reward model
D
DPO and RLHF target totally different problems — DPO post-trains the embedding head while RLHF rewrites the policy logits, and modern stacks usually run both stages independently
Válaszolj
💬
Ügyfélszolgálat