Machine learningLLMs & Text
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
CuMo, a model that integrates Co-upcycled Top-K sparsely-gated Mixture-of-experts blocks into the vision encoder and the MLP connector, improves multimodal LLMs with minimal additional activated parameters during inference, outperforming other multimodal LLMs across various benchmarks.
Featured in No. 49 on 15 May 2024 · 6 days after release · 78 citations today · published in Neural Information Processing Systems
- Released
- 9 May 2024
- First featured
- No. 49 · 15 May 2024
- Citations (Semantic Scholar)
- 78
- Influential citations
- 8
- Published in
- Neural Information Processing Systems
- Shares when featured
- 55
- Identifier
- arXiv:2405.05949
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).