ML-QuantSubscribe

Machine learningLLMs & Text

Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF

The study presents a unified approach to reinforcement learning from human feedback for large language models, offering theoretical guarantees and practical effectiveness.

Featured in No. 52 on 5 Jun 2024 · 7 days after release · 75 citations today · published in International Conference on Learning Representations

Released
29 May 2024
First featured
No. 52 · 5 Jun 2024
Citations (Semantic Scholar)
75
Influential citations
11
Published in
International Conference on Learning Representations
Shares when featured
67
Identifier
arXiv:2405.19320

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page