Skip to main navigation Skip to search Skip to main content

On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization

  • Yong Lin
  • , Skyler Seto
  • , Maartje Ter Hoeve
  • , Katherine Metcalf
  • , Barry John Theobald
  • , Xuan Wang
  • , Yizhe Zhang
  • , Chen Huang
  • , Tong Zhang

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

Reinforcement Learning from Human Feedback (RLHF) is an effective approach for aligning language models to human preferences and reducing risks in deploying them in the wild. Central to RLHF is learning a reward function for scoring human preferences. Two main approaches for learning a reward model are 1) training an EXplicit Reward Model (EXRM) as in RLHF, and 2) using an implicit reward learned from preference data through methods such as Direct Preference Optimization (DPO). Prior work has shown that the implicit reward model of DPO (denoted as DPORM) can approximate an EXRM in the limit. However, it is unclear how well DPORM empirically matches the performance of EXRM. DPORM's effectiveness directly implies the optimality of the learned policy, and also impacts preference labeling in LLM alignment methods including iterative DPO. This work studies the accuracy at distinguishing preferred and rejected answers for both DPORM and EXRM. Our findings indicate that even though DPORM fits the training dataset comparably, it generalizes less effectively than EXRM, especially when the validation datasets contain distribution shifts. Across five out-of-distribution settings, DPORM has a mean drop in accuracy of 3% and a maximum drop of 7%. These findings highlight that DPORM has limited generalization ability and substantiates the integration of an explicit reward model in iterative DPO approaches.

Original languageEnglish (US)
Title of host publicationEMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Findings of EMNLP 2024
EditorsYaser Al-Onaizan, Mohit Bansal, Yun-Nung Chen
PublisherAssociation for Computational Linguistics (ACL)
Pages16015-16026
Number of pages12
ISBN (Electronic)9798891761681
DOIs
StatePublished - 2024
Event2024 Findings of the Association for Computational Linguistics, EMNLP 2024 - Hybrid, Miami, United States
Duration: Nov 12 2024Nov 16 2024

Publication series

NameEMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Findings of EMNLP 2024

Conference

Conference2024 Findings of the Association for Computational Linguistics, EMNLP 2024
Country/TerritoryUnited States
CityHybrid, Miami
Period11/12/2411/16/24

ASJC Scopus subject areas

  • Computational Theory and Mathematics
  • Computer Science Applications
  • Information Systems
  • Linguistics and Language

Fingerprint

Dive into the research topics of 'On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization'. Together they form a unique fingerprint.

Cite this