Machine learningOther
BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts
The paper introduces BAM, a new method for training Mixture of Experts models, which optimizes the use of specialized dense models and improves performance in perplexity and downstream tasks.
Featured in No. 62 on 21 Aug 2024 · 6 days after release · 18 citations today · published in Neural Information Processing Systems
- Released
- 15 Aug 2024
- First featured
- No. 62 · 21 Aug 2024
- Citations (Semantic Scholar)
- 18
- Influential citations
- 1
- Published in
- Neural Information Processing Systems
- Shares when featured
- 47
- Identifier
- arXiv:2408.08274
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).