Прескочи на главни садржај
🔒 Режим прегледа. Првих петнаест Foundations лекција је бесплатно; ова је Pro. Покрените 7-дневни trial да откључате едитор, AI савете и остатак курса. Картица је обавезна, можете отказати у било ком тренутку у Dashboard.
Покрени 7-дневни trial →
⚡
← Kursevi
›
AI Engineering with Python
›
Модул 4 · Агент Лоопс & Воркфловс
›
Основе ДПО и РЛХФ
quiz
56 / 105
🇷🇸
SR
▼
↗
Подели
⋯
Више
+100 XP
📋
Zadatak
📖
Teorija
🤖
AI Pomoć
Zadatak
📝 **Питање:** Која је практична разлика између РЛХФ и ДПО? 📋 Изаберите тачан одговор. 💡 **Савет:** Поново прочитајте горњу теорију ако нисте сигурни.
🎯 Kviz
Pitanje
📝 **Питање:** Која је практична разлика између РЛХФ и ДПО? 📋 Изаберите тачан одговор. 💡 **Савет:** Поново прочитајте горњу теорију ако нисте сигурни.
A
DPO predates RLHF — Christiano's 2017 preference paper is the first DPO paper and OpenAI later wrapped a PPO loop around the same loss, branding the wrapper RLHF for InstructGPT
B
DPO drops the separate reward model and PPO loop — it directly trains the LLM to prefer 'chosen' over 'rejected' completions, making the pipeline simpler and cheaper while often matching RLHF quality
C
DPO needs roughly 10x more labeled preference pairs than RLHF to reach the same alignment quality, since the implicit reward signal in the contrastive loss is noisier than a fitted reward model
D
DPO and RLHF target totally different problems — DPO post-trains the embedding head while RLHF rewrites the policy logits, and modern stacks usually run both stages independently
Odgovori
💬 Diskusija
Budi prvi — postavi pitanje ili podeli savet.
Prijavi se
da bi se pridružio diskusiji. Čitanje je besplatno.
Učitavanje diskusije…