---
title: Foundation Models for Music: A Survey
url: https://www.ml-quant.com/papers/arxiv/2408.14340/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2408.14340
source_url: https://arxiv.org/abs/2408.14340
featured: 2024-09-05
citations: 62
topic: Other
---


# Foundation Models for Music: A Survey

The article discusses the influence of foundation models on the music industry, emphasizing their potential in music generation and the need for ethical research on issues like transparency and copyright.

- Source: https://arxiv.org/abs/2408.14340
- Identifier: arXiv:2408.14340
- Released: 2024-08-26
- First featured: Quant Letter No. 64 (2024-09-05): https://www.ml-quant.com/issues/2024-09-05/
- Citations (Semantic Scholar): 62
- Published in: not yet
- Topic: Other

## Related

- [Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction](https://www.ml-quant.com/papers/arxiv/2409.18124/): Visual Foundation for Dense Prediction: Lotus, a new visual foundation model, predicts annotations directly, improving inference speed and performance in zero-shot depth and normal estimation tasks.
- [A foundation model for joint segmentation, detection and recognition of biomedical objects across nine modalities](https://www.ml-quant.com/papers/arxiv/2405.12971/): Image Parsing Model: BiomedParse is a new tool for biomedical image analysis, capable of identifying 82 object types across 9 imaging modalities, enhancing accuracy in biomedical research.
- [OmniGlue: Generalizable Feature Matching with Foundation Model Guidance](https://www.ml-quant.com/papers/arxiv/2405.12979/): OmniGlue, a new image matcher that performs better on unseen image domains than previous models, is introduced in this paper.
- [SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation](https://www.ml-quant.com/papers/arxiv/2501.18564/): Visual Foundation Model for Robotic Manipulation: The new robotic manipulation system, SAM2Act, shows top-tier performance in various environments, and its memory-based version, SAM2Act+, surpasses existing methods in memory-dependent tasks.
- [StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos](https://www.ml-quant.com/papers/arxiv/2409.07447/): A novel framework converts 2D videos into immersive 3D content using foundation models, providing a practical solution for creating high-quality 3D content for devices such as Apple Vision Pro.
- [Towards Foundation Models for 3D Vision: How Close are We?](https://www.ml-quant.com/papers/arxiv/2410.10799/): A new 3D visual understanding benchmark shows that while specialized models are accurate, they are not robust, and human vision is still the most reliable 3D visual system.
