---
title: Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction
url: https://www.ml-quant.com/papers/arxiv/2409.18124/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2409.18124
source_url: https://arxiv.org/abs/2409.18124
featured: 2024-10-03
citations: 198
topic: Other
---


# Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction

Visual Foundation for Dense Prediction: Lotus, a new visual foundation model, predicts annotations directly, improving inference speed and performance in zero-shot depth and normal estimation tasks.

- Source: https://arxiv.org/abs/2409.18124
- Identifier: arXiv:2409.18124
- Released: 2024-09-26
- First featured: Quant Letter No. 68 (2024-10-03): https://www.ml-quant.com/issues/2024-10-03/
- Citations (Semantic Scholar): 198
- Published in: International Conference on Learning Representations
- Topic: Other

## Related

- [A foundation model for joint segmentation, detection and recognition of biomedical objects across nine modalities](https://www.ml-quant.com/papers/arxiv/2405.12971/): Image Parsing Model: BiomedParse is a new tool for biomedical image analysis, capable of identifying 82 object types across 9 imaging modalities, enhancing accuracy in biomedical research.
- [OmniGlue: Generalizable Feature Matching with Foundation Model Guidance](https://www.ml-quant.com/papers/arxiv/2405.12979/): OmniGlue, a new image matcher that performs better on unseen image domains than previous models, is introduced in this paper.
- [Foundation Models for Music: A Survey](https://www.ml-quant.com/papers/arxiv/2408.14340/): The article discusses the influence of foundation models on the music industry, emphasizing their potential in music generation and the need for ethical research on issues like transparency and copyright.
- [SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation](https://www.ml-quant.com/papers/arxiv/2501.18564/): Visual Foundation Model for Robotic Manipulation: The new robotic manipulation system, SAM2Act, shows top-tier performance in various environments, and its memory-based version, SAM2Act+, surpasses existing methods in memory-dependent tasks.
- [Towards Foundation Models for 3D Vision: How Close are We?](https://www.ml-quant.com/papers/arxiv/2410.10799/): A new 3D visual understanding benchmark shows that while specialized models are accurate, they are not robust, and human vision is still the most reliable 3D visual system.
- [StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos](https://www.ml-quant.com/papers/arxiv/2409.07447/): A novel framework converts 2D videos into immersive 3D content using foundation models, providing a practical solution for creating high-quality 3D content for devices such as Apple Vision Pro.
