arXiv:2602.24289cs.CVcs.LG2026-02被引 9

用两种学习策略,让视频生成更快更连贯。

Mode Seeking meets Mean Seeking for Fast Long Video Generation

论文配图:Mode Seeking meets Mean Seeking for Fast Long Video Generation
图 1 · 摘自论文原文
  • 分步训练:全局流匹配学叙事结构,局部分布匹配保细节真实
  • 生成分钟级视频仅需几步,同时提升画质和长程一致性
  • 适合需要快速生成高质量长视频的场景

将视频生成从秒级扩展到分钟级面临关键瓶颈:短视频数据丰富且高保真,但长视频数据稀缺且领域局限。为此,我们提出一种训练范式——模式寻找与均值寻找结合,通过解耦扩散变压器中的局部保真度与长期连贯性。该方法利用监督学习在长视频上训练全局流匹配头以捕捉叙事结构,同时采用局部分布匹配头,通过模式寻找反KL散度将滑动窗口对齐冻结的短视频教师模型。这一策略使学生模型能从有限长视频中学习长距离连贯性与运动规律,同时通过每段滑动窗口与教师模型对齐,继承局部真实性,实现几步完成的快速长视频生成。评估表明,该方法有效弥合了保真度与生成时长之间的差距,联合提升局部清晰度、动作自然性与长程一致性。

原文摘要 · Abstract (English)

Scaling video generation from seconds to minutes faces a critical bottleneck: while short-video data is abundant and high-fidelity, coherent long-form data is scarce and limited to narrow domains. To address this, we propose a training paradigm where Mode Seeking meets Mean Seeking, decoupling local fidelity from long-term coherence based on a unified representation via a Decoupled Diffusion Transformer. Our approach utilizes a global Flow Matching head trained via supervised learning on long videos to capture narrative structure, while simultaneously employing a local Distribution Matching head that aligns sliding windows to a frozen short-video teacher via a mode-seeking reverse-KL divergence. This strategy enables the synthesis of minute-scale videos that learns long-range coherence and motions from limited long videos via supervised flow matching, while inheriting local realism by aligning every sliding-window segment of the student to a frozen short-video teacher, resulting in a few-step fast long video generator. Evaluations show that our method effectively closes the fidelity-horizon gap by jointly improving local sharpness, motion and long-range consistency. Project website: https://primecai.github.io/mmm/.

视频生成扩散模型长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。