TY - GEN
T1 - On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
AU - Lin, Yong
AU - Seto, Skyler
AU - Ter Hoeve, Maartje
AU - Metcalf, Katherine
AU - Theobald, Barry John
AU - Wang, Xuan
AU - Zhang, Yizhe
AU - Huang, Chen
AU - Zhang, Tong
N1 - Publisher Copyright:
© 2024 Association for Computational Linguistics.
PY - 2024
Y1 - 2024
N2 - Reinforcement Learning from Human Feedback (RLHF) is an effective approach for aligning language models to human preferences and reducing risks in deploying them in the wild. Central to RLHF is learning a reward function for scoring human preferences. Two main approaches for learning a reward model are 1) training an EXplicit Reward Model (EXRM) as in RLHF, and 2) using an implicit reward learned from preference data through methods such as Direct Preference Optimization (DPO). Prior work has shown that the implicit reward model of DPO (denoted as DPORM) can approximate an EXRM in the limit. However, it is unclear how well DPORM empirically matches the performance of EXRM. DPORM's effectiveness directly implies the optimality of the learned policy, and also impacts preference labeling in LLM alignment methods including iterative DPO. This work studies the accuracy at distinguishing preferred and rejected answers for both DPORM and EXRM. Our findings indicate that even though DPORM fits the training dataset comparably, it generalizes less effectively than EXRM, especially when the validation datasets contain distribution shifts. Across five out-of-distribution settings, DPORM has a mean drop in accuracy of 3% and a maximum drop of 7%. These findings highlight that DPORM has limited generalization ability and substantiates the integration of an explicit reward model in iterative DPO approaches.
AB - Reinforcement Learning from Human Feedback (RLHF) is an effective approach for aligning language models to human preferences and reducing risks in deploying them in the wild. Central to RLHF is learning a reward function for scoring human preferences. Two main approaches for learning a reward model are 1) training an EXplicit Reward Model (EXRM) as in RLHF, and 2) using an implicit reward learned from preference data through methods such as Direct Preference Optimization (DPO). Prior work has shown that the implicit reward model of DPO (denoted as DPORM) can approximate an EXRM in the limit. However, it is unclear how well DPORM empirically matches the performance of EXRM. DPORM's effectiveness directly implies the optimality of the learned policy, and also impacts preference labeling in LLM alignment methods including iterative DPO. This work studies the accuracy at distinguishing preferred and rejected answers for both DPORM and EXRM. Our findings indicate that even though DPORM fits the training dataset comparably, it generalizes less effectively than EXRM, especially when the validation datasets contain distribution shifts. Across five out-of-distribution settings, DPORM has a mean drop in accuracy of 3% and a maximum drop of 7%. These findings highlight that DPORM has limited generalization ability and substantiates the integration of an explicit reward model in iterative DPO approaches.
UR - https://www.scopus.com/pages/publications/85217622169
UR - https://www.scopus.com/pages/publications/85217622169#tab=citedBy
U2 - 10.18653/v1/2024.findings-emnlp.940
DO - 10.18653/v1/2024.findings-emnlp.940
M3 - Conference contribution
AN - SCOPUS:85217622169
T3 - EMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Findings of EMNLP 2024
SP - 16015
EP - 16026
BT - EMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Findings of EMNLP 2024
A2 - Al-Onaizan, Yaser
A2 - Bansal, Mohit
A2 - Chen, Yun-Nung
PB - Association for Computational Linguistics (ACL)
T2 - 2024 Findings of the Association for Computational Linguistics, EMNLP 2024
Y2 - 12 November 2024 through 16 November 2024
ER -