ML-QuantSubscribe

Machine learningLLMs & Text

Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment

A study has found that language models aligned with human-annotated preference data are preferred by humans in over 70% of cases, even without language-specific data for supervised finetuning.

Featured in No. 70 on 17 Oct 2024 · · 32 citations today · published in Conference on Empirical Methods in Natural Language Processing

Released
18 Apr 2024
First featured
No. 70 · 17 Oct 2024
Citations (Semantic Scholar)
32
Influential citations
1
Published in
Conference on Empirical Methods in Natural Language Processing
Shares when featured
119
Identifier
arXiv:2404.12318

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page