ML-QuantSubscribe

arXivLLMs & Text

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models

The paper introduces a new method for training large language models using reinforcement learning feedback, which makes models more beneficial and less harmful by adjusting sensitivity to reward values.

Featured in No. 82 on 15 Jan 2025 · 7 days after release · 2 citations today

Released
8 Jan 2025
First featured
No. 82 · 15 Jan 2025
Citations (Semantic Scholar)
2
Influential citations
0
Published in
Not yet, as far as Semantic Scholar knows
Shares when featured
17
Identifier
arXiv:2501.06248

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page