---
title: An Empirical Study of $\mu$P Learning Rate Transfer
url: https://www.ml-quant.com/papers/arxiv/2404.05728/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2404.05728
source_url: https://arxiv.org/abs/2404.05728
featured: 2024-04-17
citations: 7
topic: ML & AI Methods
---


# An Empirical Study of $\mu$P Learning Rate Transfer

Neural Network Scaling Rules: A study has found that the μ-Parameterization (μP) is generally effective in determining the best learning rates for large neural network models, although it doesn't work in all situations.

- Source: https://arxiv.org/abs/2404.05728
- Identifier: arXiv:2404.05728
- Released: 2024-04-08
- First featured: Quant Letter No. 45 (2024-04-17): https://www.ml-quant.com/issues/2024-04-17/
- Citations (Semantic Scholar): 7
- Published in: not yet
- Topic: ML & AI Methods

## Related

- [Mamba: Linear-Time Sequence Modeling with Selective State Spaces](https://www.ml-quant.com/papers/arxiv/2312.00752/): Sequence Modeling: Mamba, a neural network architecture that doesn't use attention or MLP blocks, provides faster inference and better performance in language, audio, and genomics than Transformers.
- [Graph Mamba: Towards Learning on Graphs with State Space Models](https://www.ml-quant.com/papers/arxiv/2402.08678/): Graph Mamba Networks, a new type of Graph Neural Networks, have been introduced, which achieve excellent performance in various benchmark datasets despite lower computational cost.
- [Edge Directionality Improves Learning on Heterophilic Graphs](https://www.ml-quant.com/papers/arxiv/2305.10498/): The study presents Directed Graph Neural Network (Dir-GNN), a new deep learning framework for directed graphs that surpasses traditional models in heterophilic benchmarks.
- [Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks](https://www.ml-quant.com/papers/arxiv/2310.02244/): Deep Residual Network Feature Learning: The research explores depthwise parametrizations in deep residual networks, pinpointing Depth-$\mu$P as the best parametrization for maximizing feature learning and diversity, but notes its limitations in deeper networks.
- [Modular Duality in Deep Learning](https://www.ml-quant.com/papers/arxiv/2410.21265/): The article presents a new theory of modular dualization for general neural networks, providing a theoretical basis for fast and scalable training algorithms, potentially leading to a new generation of optimizers for neural architectures.
- [Enhancing path-integral approximation for non-linear diffusion with neural network](https://www.ml-quant.com/papers/arxiv/2404.08903/): The paper improves the pricing of fixed income instruments within the Black-Karasinski model using neural networks, showing better results for multiple calibrations over extended periods.
