通过时间扰动提升视频生成的连贯性与多样性
Temporal Regularization Makes Your Video Generator Stronger
- 在数据层面引入可控时间扰动,无需修改模型结构
- 在UCF-101和VBench上显著提升多类模型的时间一致性
- 适合追求视频生成质量优化的研究者与开发者
时间质量是视频生成的关键,直接影响运动连续性与动态真实性。本文首次探索视频生成中的时间增强,提出FluxFlow策略以提升时间质量。该方法在数据层施加受控时间扰动,无需架构改动。在UCF-101和VBench基准上的大量实验表明,FluxFlow显著提升U-Net、DiT及基于AR的多种模型的时间连贯性与多样性,同时保持空间保真度。结果表明,时间增强是一种简单而有效的视频生成质量提升方法。
原文摘要 · Abstract (English)
Temporal quality is a critical aspect of video generation, as it ensures consistent motion and realistic dynamics across frames. However, achieving high temporal coherence and diversity remains challenging. In this work, we explore temporal augmentation in video generation for the first time, and introduce FluxFlow for initial investigation, a strategy designed to enhance temporal quality. Operating at the data level, FluxFlow applies controlled temporal perturbations without requiring architectural modifications. Extensive experiments on UCF-101 and VBench benchmarks demonstrate that FluxFlow significantly improves temporal coherence and diversity across various video generation models, including U-Net, DiT, and AR-based architectures, while preserving spatial fidelity. These findings highlight the potential of temporal augmentation as a simple yet effective approach to advancing video generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。