Machine learningLLMs & Text
Dynamic Rewarding with Prompt Optimization Enables Tuning-free Self-Alignment of Language Models
The paper presents a new tuning-free method called Dynamic Rewarding with Prompt Optimization (DRPO) for self-aligning Large Language Models (LLMs), improving alignment performance without extra training or human intervention.
Featured in No. 75 on 20 Nov 2024 · 7 days after release · 20 citations today · published in Conference on Empirical Methods in Natural Language Processing
- Released
- 13 Nov 2024
- First featured
- No. 75 · 20 Nov 2024
- Citations (Semantic Scholar)
- 20
- Influential citations
- 1
- Published in
- Conference on Empirical Methods in Natural Language Processing
- Shares when featured
- 27
- Identifier
- arXiv:2411.08733
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).