Machine learningML & AI Methods
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
A novel method to prevent reward hacking in AI systems uses state occupancy measure instead of action distribution, effectively avoiding significant drops in true reward.
Featured in No. 39 on 6 Mar 2024 · 1 day after release · 55 citations today · published in International Conference on Learning Representations
- Released
- 5 Mar 2024
- First featured
- No. 39 · 6 Mar 2024
- Citations (Semantic Scholar)
- 55
- Influential citations
- 3
- Published in
- International Conference on Learning Representations
- Shares when featured
- 10
- Identifier
- arXiv:2403.03185
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).