---
title: Towards Foundation Models for 3D Vision: How Close are We?
url: https://www.ml-quant.com/papers/arxiv/2410.10799/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2410.10799
source_url: https://arxiv.org/pdf/2410.10799
featured: 2024-10-17
citations: 14
topic: Other
---


# Towards Foundation Models for 3D Vision: How Close are We?

A new 3D visual understanding benchmark shows that while specialized models are accurate, they are not robust, and human vision is still the most reliable 3D visual system.

- Source: https://arxiv.org/pdf/2410.10799
- Identifier: arXiv:2410.10799
- Released: 2024-10-14
- First featured: Quant Letter No. 70 (2024-10-17): https://www.ml-quant.com/issues/2024-10-17/
- Citations (Semantic Scholar): 14
- Published in: 2025 International Conference on 3D Vision (3DV)
- Topic: Other

## Related

- [Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction](https://www.ml-quant.com/papers/arxiv/2409.18124/): Visual Foundation for Dense Prediction: Lotus, a new visual foundation model, predicts annotations directly, improving inference speed and performance in zero-shot depth and normal estimation tasks.
- [A foundation model for joint segmentation, detection and recognition of biomedical objects across nine modalities](https://www.ml-quant.com/papers/arxiv/2405.12971/): Image Parsing Model: BiomedParse is a new tool for biomedical image analysis, capable of identifying 82 object types across 9 imaging modalities, enhancing accuracy in biomedical research.
- [OmniGlue: Generalizable Feature Matching with Foundation Model Guidance](https://www.ml-quant.com/papers/arxiv/2405.12979/): OmniGlue, a new image matcher that performs better on unseen image domains than previous models, is introduced in this paper.
- [Foundation Models for Music: A Survey](https://www.ml-quant.com/papers/arxiv/2408.14340/): The article discusses the influence of foundation models on the music industry, emphasizing their potential in music generation and the need for ethical research on issues like transparency and copyright.
- [SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation](https://www.ml-quant.com/papers/arxiv/2501.18564/): Visual Foundation Model for Robotic Manipulation: The new robotic manipulation system, SAM2Act, shows top-tier performance in various environments, and its memory-based version, SAM2Act+, surpasses existing methods in memory-dependent tasks.
- [StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos](https://www.ml-quant.com/papers/arxiv/2409.07447/): A novel framework converts 2D videos into immersive 3D content using foundation models, providing a practical solution for creating high-quality 3D content for devices such as Apple Vision Pro.
