Machine learningLLMs & Text
From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
The research explores Direct Preference Optimization (DPO) in Reinforcement Learning From Human Feedback (RLHF), showing its ability to assign credit and its similarity to search-based algorithms in language generation.
Featured in No. 46 on 24 Apr 2024 · 6 days after release · 273 citations today
- Released
- 18 Apr 2024
- First featured
- No. 46 · 24 Apr 2024
- Citations (Semantic Scholar)
- 273
- Influential citations
- 27
- Published in
- Not yet, as far as Semantic Scholar knows
- Shares when featured
- 107
- Identifier
- arXiv:2404.12358
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).