---
title: T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation
url: https://www.ml-quant.com/papers/arxiv/2407.14505/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2407.14505
source_url: https://arxiv.org/abs/2407.14505
featured: 2025-01-23
citations: 183
topic: ML & AI Methods
---


# T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation

Text-to-Video Benchmark: TV-CompBench, a new benchmark for evaluating text-to-video generative models, shows that current models struggle with composing various elements into a video.

- Source: https://arxiv.org/abs/2407.14505
- Identifier: arXiv:2407.14505
- Released: 2024-07-19
- First featured: Quant Letter No. 83 (2025-01-23): https://www.ml-quant.com/issues/2025-01-23/
- Citations (Semantic Scholar): 183
- Published in: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
- Topic: ML & AI Methods

## Related

- [CAT3D: Create Anything in 3D with Multi-View Diffusion Models](https://www.ml-quant.com/papers/arxiv/2405.10314/): Multi-View Diffusion Models: CAT3D is a novel technique for generating 3D scenes from any number of images, surpassing existing methods in speed and efficiency.
- [Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models](https://www.ml-quant.com/papers/arxiv/2501.01423/): The paper proposes a new model, VA-VAE, that aligns the latent space with pre-trained vision foundation models, enabling faster convergence of Diffusion Transformers in high-dimensional latent spaces and achieving top performance on ImageNet 256x256 generation.
- [Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps](https://www.ml-quant.com/papers/arxiv/2501.09732/): The research shows that increasing computation during inference-time can enhance the quality of samples produced by diffusion models, especially in image generation.
- [Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control](https://www.ml-quant.com/papers/arxiv/2409.08861/): The study presents Adjoint Matching, a new algorithm that enhances dynamical generative models by refining reward fine-tuning, leading to improved consistency, realism, and adaptability to unseen human preference reward models.
- [ShieldGemma: Generative AI Content Moderation Based on Gemma](https://www.ml-quant.com/papers/arxiv/2407.21772/): ShieldGemma is a safety content moderation model that excels in predicting safety risks such as explicit content and hate speech, surpassing models like LlamaGuard and WildCard.
- [Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens](https://www.ml-quant.com/papers/arxiv/2410.13863/): The study explores text-to-image generation, finding continuous token-based models offer superior visual quality and random-order models score higher on the GenEval benchmark, leading to a new model, Fluid.
