---
title: Improving the Training of Rectified Flows
url: https://www.ml-quant.com/papers/arxiv/2405.20320/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2405.20320
source_url: https://arxiv.org/abs/2405.20320
featured: 2024-06-05
citations: 118
topic: ML & AI Methods
---


# Improving the Training of Rectified Flows

The paper suggests improved methods for training rectified flows in diffusion models for image and video generation, surpassing current distillation methods in low function evaluation settings.

- Source: https://arxiv.org/abs/2405.20320
- Identifier: arXiv:2405.20320
- Released: 2024-05-30
- First featured: Quant Letter No. 52 (2024-06-05): https://www.ml-quant.com/issues/2024-06-05/
- Citations (Semantic Scholar): 118
- Published in: Neural Information Processing Systems
- Topic: ML & AI Methods

## Related

- [CAT3D: Create Anything in 3D with Multi-View Diffusion Models](https://www.ml-quant.com/papers/arxiv/2405.10314/): Multi-View Diffusion Models: CAT3D is a novel technique for generating 3D scenes from any number of images, surpassing existing methods in speed and efficiency.
- [Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models](https://www.ml-quant.com/papers/arxiv/2501.01423/): The paper proposes a new model, VA-VAE, that aligns the latent space with pre-trained vision foundation models, enabling faster convergence of Diffusion Transformers in high-dimensional latent spaces and achieving top performance on ImageNet 256x256 generation.
- [Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps](https://www.ml-quant.com/papers/arxiv/2501.09732/): The research shows that increasing computation during inference-time can enhance the quality of samples produced by diffusion models, especially in image generation.
- [ShieldGemma: Generative AI Content Moderation Based on Gemma](https://www.ml-quant.com/papers/arxiv/2407.21772/): ShieldGemma is a safety content moderation model that excels in predicting safety risks such as explicit content and hate speech, surpassing models like LlamaGuard and WildCard.
- [Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control](https://www.ml-quant.com/papers/arxiv/2409.08861/): The study presents Adjoint Matching, a new algorithm that enhances dynamical generative models by refining reward fine-tuning, leading to improved consistency, realism, and adaptability to unseen human preference reward models.
- [Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion](https://www.ml-quant.com/papers/arxiv/2308.12469/): A new method using self-attention layers in stable diffusion models achieves superior zero-shot segmentation without annotations, outperforming previous methods on the COCO-Stuff-27 dataset.
