arXiv:2502.06764cs.LGcs.CV2025-02ICML被引 158

让视频生成更连贯,用历史帧动态引导扩散模型。

History-Guided Video Diffusion

论文配图:History-Guided Video Diffusion
图 1 · 摘自论文原文
  • 提出DFoT架构与历史引导机制,支持可变长度历史帧条件生成。
  • 历史引导使视频质量与时间一致性显著提升,长视频生成更稳定。
  • 适合需要高质量、长时序一致视频生成的研究者与开发者。

Classifier-free guidance(CFG)是提升扩散模型条件生成能力的关键技术,能增强控制精度并提高样本质量。将其拓展至视频生成——即基于可变数量的上下文帧(统称历史)生成视频——面临两大挑战:现有架构仅支持固定大小的条件输入,且经验发现CFG式的历史帧丢弃效果不佳。为此,我们提出Diffusion Forcing Transformer(DFoT),一种支持灵活历史帧数的视频扩散架构及理论严谨的训练目标。进一步提出历史引导(History Guidance),一类由DFoT支持的独特引导方法。结果显示,最简单的原始历史引导已显著提升视频生成质量与时间一致性;更先进的跨时频历史引导则进一步增强运动动态,实现分布外历史的组合泛化,并可稳定生成极长视频。项目主页:https://boyuan.space/history-guidance

原文摘要 · Abstract (English)

Classifier-free guidance (CFG) is a key technique for improving conditional generation in diffusion models, enabling more accurate control while enhancing sample quality. It is natural to extend this technique to video diffusion, which generates video conditioned on a variable number of context frames, collectively referred to as history. However, we find two key challenges to guiding with variable-length history: architectures that only support fixed-size conditioning, and the empirical observation that CFG-style history dropout performs poorly. To address this, we propose the Diffusion Forcing Transformer (DFoT), a video diffusion architecture and theoretically grounded training objective that jointly enable conditioning on a flexible number of history frames. We then introduce History Guidance, a family of guidance methods uniquely enabled by DFoT. We show that its simplest form, vanilla history guidance, already significantly improves video generation quality and temporal consistency. A more advanced method, history guidance across time and frequency further enhances motion dynamics, enables compositional generalization to out-of-distribution history, and can stably roll out extremely long videos. Project website: https://boyuan.space/history-guidance

视频生成扩散模型历史引导时序一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。