---
title: Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos
url: https://www.ml-quant.com/papers/arxiv/2412.09621/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2412.09621
source_url: https://arxiv.org/abs/2412.09621
featured: 2024-12-18
citations: 94
topic: ML & AI Methods
---


# Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos

Learning Motion from Videos: The authors have developed a system that mines high-quality 4D reconstructions from internet videos, allowing for the prediction of structure and 3D motion from real-world image pairs.

- Source: https://arxiv.org/abs/2412.09621
- Identifier: arXiv:2412.09621
- Released: 2024-12-12
- First featured: Quant Letter No. 79 (2024-12-18): https://www.ml-quant.com/issues/2024-12-18/
- Citations (Semantic Scholar): 94
- Published in: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
- Topic: ML & AI Methods

## Related

- [Mamba: Linear-Time Sequence Modeling with Selective State Spaces](https://www.ml-quant.com/papers/arxiv/2312.00752/): Sequence Modeling: Mamba, a neural network architecture that doesn't use attention or MLP blocks, provides faster inference and better performance in language, audio, and genomics than Transformers.
- [Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality](https://www.ml-quant.com/papers/arxiv/2405.21060/): The research identifies a link between state-space models and Transformers in deep learning, leading to the creation of a faster language modeling architecture, Mamba-2.
- [Octo: An Open-Source Generalist Robot Policy](https://www.ml-quant.com/papers/arxiv/2405.12213/): Octo is a large transformer-based policy for robotic manipulation, trained on a vast dataset, that can be instructed via language or images and adapted to new domains.
- [Mastering Diverse Domains through World Models](https://www.ml-quant.com/papers/arxiv/2301.04104/): Algorithm Mastery: DreamerV3, a universal algorithm, excels in over 150 varied tasks, including diamond collection in Minecraft without human input, expanding the scope of reinforcement learning.
- [SimPO: Simple Preference Optimization with a Reference-Free Reward](https://www.ml-quant.com/papers/arxiv/2405.14734/): Simple Preference Optimization: SimPO improves reinforcement learning from human feedback by using the average log probability of a sequence as the implicit reward, enhancing training stability and computational efficiency.
- [Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction](https://www.ml-quant.com/papers/arxiv/2404.02905/): The article discusses Visual AutoRegressive modeling (VAR), a new image learning method that outperforms diffusion transformers in terms of speed, image quality, and scalability.
