---
title: GPS as a Control Signal for Image Generation
url: https://www.ml-quant.com/papers/arxiv/2501.12390/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2501.12390
source_url: https://arxiv.org/abs/2501.12390
featured: 2025-01-23
citations: 8
topic: Other
---


# GPS as a Control Signal for Image Generation

The research uses GPS tags in photo metadata to train models that generate images based on location, improving the estimated 3D structure and capturing the unique appearance of different locations.

- Source: https://arxiv.org/abs/2501.12390
- Identifier: arXiv:2501.12390
- Released: 2025-01-21
- First featured: Quant Letter No. 83 (2025-01-23): https://www.ml-quant.com/issues/2025-01-23/
- Citations (Semantic Scholar): 8
- Published in: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
- Topic: Other

## Related

- [Depth Anything V2](https://www.ml-quant.com/papers/arxiv/2406.09414/): Depth Anything V2 is a new model for monocular depth estimation, using synthetic and large-scale pseudo-labeled real images for faster, more accurate results and setting a new evaluation benchmark.
- [MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark](https://www.ml-quant.com/papers/arxiv/2406.01574/): MMLU-Pro, an improved dataset, expands the Massive Multitask Language Understanding benchmark by adding tougher questions and more choices, serving as a better benchmark to monitor progress in the field.
- [Qwen2.5-Coder Technical Report](https://www.ml-quant.com/papers/arxiv/2409.12186/): The report unveils the Qwen2.5-Coder series, an improvement from its predecessor, showcasing remarkable code generation abilities and achieving top-tier performance in various code-related tasks.
- [Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives](https://www.ml-quant.com/papers/arxiv/2311.18259/): Understanding Human Activity: The paper presents Ego-Exo4D, a large-scale video dataset and benchmark challenge featuring human activities from various perspectives, aimed at improving first-person video understanding.
- [Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian Splatting](https://www.ml-quant.com/papers/arxiv/2310.10642/): The 4DGS model is introduced, capable of reconstructing dynamic 3D scenes from 2D images and generating diverse views over time, providing real-time rendering efficiency.
- [Continuous 3D Perception Model with Persistent State](https://www.ml-quant.com/papers/arxiv/2501.12387/): The paper presents CUT3R, a unified framework that uses a recurrent model to generate metric-scale pointmaps from a stream of images, enabling dense scene reconstruction that updates with new images.
