ML-QuantSubscribe

Machine learningML & AI Methods

Dataset Reset Policy Optimization for RLHF

The DR-PO algorithm enhances Reinforcement Learning by incorporating offline preference data into online policy training, outperforming other techniques in summarization and the Anthropic Helpful Harmful dataset.

Featured in No. 45 on 17 Apr 2024 · 5 days after release · 44 citations today

Released
12 Apr 2024
First featured
No. 45 · 17 Apr 2024
Citations (Semantic Scholar)
44
Influential citations
7
Published in
Not yet, as far as Semantic Scholar knows
Shares when featured
23
Identifier
arXiv:2404.08495

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page