用视频帧时序信息提升图像扩散模型训练速度与多样性
Fitting Image Diffusion Models on Video Datasets
- 利用连续视频帧的时序依赖性改进扩散模型训练
- 训练速度提升2倍以上,FID更低,生成更丰富
- 无需修改结构,适合快速集成到现有扩散模型中
图像扩散模型通常在独立采样的静态图像上训练,这种设计本质上无法捕捉时间信息,导致收敛慢、分布覆盖有限、泛化能力差。本文提出一种简单有效的训练策略,利用连续视频帧中的时序归纳偏置来改进扩散模型训练。该方法无需架构修改,可无缝集成至标准训练流程。在手物交互数据集HandCo上评估显示,模型收敛速度提升超过2倍,训练与验证分布上的FID均更低,且生成多样性增强,能更好捕捉细微的指关节运动差异。优化分析表明,该正则化降低了梯度方差,从而加速收敛。
原文摘要 · Abstract (English)
Image diffusion models are trained on independently sampled static images. While this is the bedrock task protocol in generative modeling, capturing the temporal world through the lens of static snapshots is information-deficient by design. This limitation leads to slower convergence, limited distributional coverage, and reduced generalization. In this work, we propose a simple and effective training strategy that leverages the temporal inductive bias present in continuous video frames to improve diffusion training. Notably, the proposed method requires no architectural modification and can be seamlessly integrated into standard diffusion training pipelines. We evaluate our method on the HandCo dataset, where hand-object interactions exhibit dense temporal coherence and subtle variations in finger articulation often result in semantically distinct motions. Empirically, our method accelerates convergence by over 2$\text{x}$ faster and achieves lower FID on both training and validation distributions. It also improves generative diversity by encouraging the model to capture meaningful temporal variations. We further provide an optimization analysis showing that our regularization reduces the gradient variance, which contributes to faster convergence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。