Machine learningML & AI Methods
Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order Perspective
The research explores the performance of adaptive gradient optimizers without the square root, showing they maintain performance on transformers and improve on convolutional architectures.
Featured in No. 64 on 5 Sep 2024 · · 27 citations today · published in International Conference on Machine Learning
- Released
- 5 Feb 2024
- First featured
- No. 64 · 5 Sep 2024
- Citations (Semantic Scholar)
- 27
- Influential citations
- 2
- Published in
- International Conference on Machine Learning
- Shares when featured
- 45
- Identifier
- arXiv:2402.03496
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).