ML-QuantSubscribe

Machine learningLLMs & Text

Self-Play Preference Optimization for Language Model Alignment

The article suggests a self-play-based method, SPPO, for language model alignment, which can effectively enhance the likelihood of the selected response and reduce that of the discarded one.

Featured in No. 51 on 28 May 2024 · 27 days after release · 278 citations today · published in International Conference on Learning Representations

Released
1 May 2024
First featured
No. 51 · 28 May 2024
Citations (Semantic Scholar)
278
Influential citations
41
Published in
International Conference on Learning Representations
Shares when featured
309
Identifier
arXiv:2405.00675

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page