让视频每帧独立降噪,提升生成质量与灵活性
Redefining Temporal Modeling in Video Diffusion: The Vectorized Timestep Approach
- 为每帧设计独立噪声调度,突破传统统一时间步限制
- 在图像转视频、长视频生成等任务中显著优于现有方法
- 适合需要精细时序控制的视频生成研究者
扩散模型已革新图像生成,其视频生成扩展也展现出潜力。然而,当前视频扩散模型(VDM)依赖全局标量时间步,难以捕捉复杂时序依赖,尤其在图像转视频等任务中表现受限。为此,我们提出帧感知视频扩散模型(FVDM),引入新型向量化时间步(VTV),使每帧可遵循独立噪声调度,增强对细粒度时序关系的建模能力。FVDM在标准视频生成、图像转视频、视频插值和长视频合成等任务中均表现出色。通过多种VTV配置,模型在生成质量上超越现有方法,有效缓解微调中的灾难性遗忘问题,并提升零样本方法的泛化能力。实证评估表明,FVDM在视频生成质量及任务扩展性方面均优于当前最优方法。该工作为视频生成树立了新范式,对生成建模与多媒体应用具有重要意义。
原文摘要 · Abstract (English)
Diffusion models have revolutionized image generation, and their extension to video generation has shown promise. However, current video diffusion models~(VDMs) rely on a scalar timestep variable applied at the clip level, which limits their ability to model complex temporal dependencies needed for various tasks like image-to-video generation. To address this limitation, we propose a frame-aware video diffusion model~(FVDM), which introduces a novel vectorized timestep variable~(VTV). Unlike conventional VDMs, our approach allows each frame to follow an independent noise schedule, enhancing the model's capacity to capture fine-grained temporal dependencies. FVDM's flexibility is demonstrated across multiple tasks, including standard video generation, image-to-video generation, video interpolation, and long video synthesis. Through a diverse set of VTV configurations, we achieve superior quality in generated videos, overcoming challenges such as catastrophic forgetting during fine-tuning and limited generalizability in zero-shot methods.Our empirical evaluations show that FVDM outperforms state-of-the-art methods in video generation quality, while also excelling in extended tasks. By addressing fundamental shortcomings in existing VDMs, FVDM sets a new paradigm in video synthesis, offering a robust framework with significant implications for generative modeling and multimedia applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。