Preskoči na glavni sadržaj
🔒 Način pregleda. Prvih petnaest Foundations lekcija je besplatno; ova je Pro. Pokrenite 7-dnevni trial da otključate editor, AI savjete i ostatak programa. Kartica obavezna, otkažite bilo kada u Dashboard.
Pokreni 7-dnevni trial →
⚡
← Kolegiji
›
AI Engineering with Python
›
Modul 4 · Agentske petlje i tijek rada
›
DPO i RLHF osnove
quiz
56 / 105
🇭🇷
HR
▼
↗
Podijeli
⋯
Više
+100 XP
📋
Zadatak
📖
Teorija
🤖
AI Pomoć
Zadatak
🌐 shown in EN
📝 **Question:** What's the practical difference between RLHF and DPO? 📋 Pick the right answer. 💡 **Hint:** Re-read the theory above if unsure.
🎯 Kviz
Pitanje
📝 **Question:** What's the practical difference between RLHF and DPO? 📋 Pick the right answer. 💡 **Hint:** Re-read the theory above if unsure.
A
DPO predates RLHF — Christiano's 2017 preference paper is the first DPO paper and OpenAI later wrapped a PPO loop around the same loss, branding the wrapper RLHF for InstructGPT
B
DPO drops the separate reward model and PPO loop — it directly trains the LLM to prefer 'chosen' over 'rejected' completions, making the pipeline simpler and cheaper while often matching RLHF quality
C
DPO needs roughly 10x more labeled preference pairs than RLHF to reach the same alignment quality, since the implicit reward signal in the contrastive loss is noisier than a fitted reward model
D
DPO and RLHF target totally different problems — DPO post-trains the embedding head while RLHF rewrites the policy logits, and modern stacks usually run both stages independently
Odgovori
💬
Kontaktiraj podršku