ML-QuantSubscribe

Machine learningLLMs & Text

Understanding Reference Policies in Direct Preference Optimization

The research reveals that Direct Preference Optimization in large language models is sensitive to the KL divergence constraint and performs better with stronger reference policies.

Featured in No. 63 on 28 Aug 2024 · 41 days after release · 23 citations today · published in North American Chapter of the Association for Computational Linguistics

Released
18 Jul 2024
First featured
No. 63 · 28 Aug 2024
Citations (Semantic Scholar)
23
Influential citations
1
Published in
North American Chapter of the Association for Computational Linguistics
Shares when featured
18
Identifier
arXiv:2407.13709

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page