Machine learningLLMs & Text
Understanding Reference Policies in Direct Preference Optimization
The research reveals that Direct Preference Optimization in large language models is sensitive to the KL divergence constraint and performs better with stronger reference policies.
Featured in No. 63 on 28 Aug 2024 · 41 days after release · 23 citations today · published in North American Chapter of the Association for Computational Linguistics
- Released
- 18 Jul 2024
- First featured
- No. 63 · 28 Aug 2024
- Citations (Semantic Scholar)
- 23
- Influential citations
- 1
- Published in
- North American Chapter of the Association for Computational Linguistics
- Shares when featured
- 18
- Identifier
- arXiv:2407.13709
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).