arXiv:2603.25527cs.CV2026-03中稿 · CVPR被引 3

通过时间步选择训练,用质量不平衡数据生成更优视频。

Beyond the Golden Data: Resolving the Motion-Vision Quality Dilemma via Timestep Selective Training

  • 根据模型学习阶段动态选择不同时间步采样数据,解耦视觉与运动质量。
  • 仅用质量不平衡数据训练,性能超越使用优质数据的传统方法。
  • 适合视频生成研究者,尤其关注数据效率与训练优化的团队。

近期视频生成模型取得显著进展,但高度依赖兼具高视觉质量和高运动质量的黄金数据。本文揭示视频数据整理中的核心挑战——运动-视觉质量困境:视觉质量与运动强度存在负相关,难以同时兼顾。我们分析视频扩散模型的分层学习动态,并对质量退化的样本进行梯度分析,发现质量失衡数据在特定时间步产生的梯度与黄金数据相似。基于此,提出训练过程中的时间步选择新范式。设计时间感知质量解耦(TQD)机制,调整数据采样分布以匹配模型学习进程:运动丰富的数据倾向在较高时间步采样,而高视觉质量数据则优先在较低时间步采样。大量实验表明,仅使用分离的质量不平衡数据训练,即可实现超越传统优质数据训练的性能,挑战了完美数据的必要性。此外,该方法在高质量数据上也提升模型表现,适用于多种数据场景。

原文摘要 · Abstract (English)

Recent advances in video generation models have achieved impressive results. However, these models heavily rely on the use of high-quality data that combines both high visual quality and high motion quality. In this paper, we identify a key challenge in video data curation: the Motion-Vision Quality Dilemma. We discovered that visual quality and motion intensity inherently exhibit a negative correlation, making it hard to obtain golden data that excels in both aspects. To address this challenge, we first examine the hierarchical learning dynamics of video diffusion models and conduct gradient-based analysis on quality-degraded samples. We discover that quality-imbalanced data can produce gradients similar to golden data at appropriate timesteps. Based on this, we introduce the novel concept of Timestep selection in Training Process. We propose Timestep-aware Quality Decoupling (TQD), which modifies the data sampling distribution to better match the model's learning process. For certain types of data, the sampling distribution is skewed toward higher timesteps for motion-rich data, while high visual quality data is more likely to be sampled during lower timesteps. Through extensive experiments, we demonstrate that TQD enables training exclusively on separated imbalanced data to achieve performance surpassing conventional training with better data, challenging the necessity of perfect data in video generation. Moreover, our method also boosts model performance when trained on high-quality data, showcasing its effectiveness across different data scenarios.

视频生成扩散模型数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。