---
title: Depth Any Video with Scalable Synthetic Data
url: https://www.ml-quant.com/papers/arxiv/2410.10815/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2410.10815
source_url: https://arxiv.org/abs/2410.10815
featured: 2024-10-17
citations: 66
topic: ML & AI Methods
---


# Depth Any Video with Scalable Synthetic Data

The article presents Depth Any Video, a new model that uses synthetic data and video diffusion models to estimate video depth more accurately and consistently than previous models.

- Source: https://arxiv.org/abs/2410.10815
- Identifier: arXiv:2410.10815
- Released: 2024-10-14
- First featured: Quant Letter No. 70 (2024-10-17): https://www.ml-quant.com/issues/2024-10-17/
- Citations (Semantic Scholar): 66
- Published in: International Conference on Learning Representations
- Topic: ML & AI Methods

## Related

- [Generative AI for Synthetic Data Creation](https://www.ml-quant.com/papers/ssrn/5268010/): The paper discusses the use of Generative AI models for synthetic data generation, and how synthetic data can enhance model performance and facilitate privacy-preserving data sharing.
- [CAT3D: Create Anything in 3D with Multi-View Diffusion Models](https://www.ml-quant.com/papers/arxiv/2405.10314/): Multi-View Diffusion Models: CAT3D is a novel technique for generating 3D scenes from any number of images, surpassing existing methods in speed and efficiency.
- [Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models](https://www.ml-quant.com/papers/arxiv/2501.01423/): The paper proposes a new model, VA-VAE, that aligns the latent space with pre-trained vision foundation models, enabling faster convergence of Diffusion Transformers in high-dimensional latent spaces and achieving top performance on ImageNet 256x256 generation.
- [Machine Learning for Synthetic Data Generation: a Review](https://www.ml-quant.com/papers/arxiv/2302.04062/): A Review: The article reviews machine learning models for creating synthetic data, discussing their uses, methods, privacy issues, fairness, and future research opportunities in fields like computer vision, speech, natural language processing, healthcare, and business.
- [Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps](https://www.ml-quant.com/papers/arxiv/2501.09732/): The research shows that increasing computation during inference-time can enhance the quality of samples produced by diffusion models, especially in image generation.
- [ShieldGemma: Generative AI Content Moderation Based on Gemma](https://www.ml-quant.com/papers/arxiv/2407.21772/): ShieldGemma is a safety content moderation model that excels in predicting safety risks such as explicit content and hate speech, surpassing models like LlamaGuard and WildCard.
