arXiv:2502.11234cs.CV2025-02被引 12

用离散流生成超长视频,效率高且灵活。

MaskFlow: Discrete Flows For Flexible and Efficient Long Video Generation

  • 用帧级掩码训练,让模型跳过已生成部分
  • 可生成长达训练序列10倍的视频,质量媲美顶尖方法
  • 支持快速采样,适配多种模型和生成模式

长视频高质量生成仍面临时空动态复杂性和硬件限制的挑战。本文提出MaskFlow,一种统一的视频生成框架,结合离散表示与流匹配技术,实现高效长视频生成。通过训练时采用帧级掩码策略,模型基于已生成的未掩码帧进行条件生成,可生成长度达训练序列十倍的视频。该方法通过支持快速掩码生成模型(MGM)式采样,显著提升效率,同时兼容全自回归与全序列生成模式。我们在FaceForensics(FFS)和Deepmind Lab(DMLab)数据集上验证了方法质量,报告的弗雷切特视频距离(FVD)达到当前最优水平。此外,我们分析了采样效率,证明MaskFlow可在无需重新训练的情况下应用于依赖时间步与不依赖时间步的模型。

原文摘要 · Abstract (English)

Generating long, high-quality videos remains a challenge due to the complex interplay of spatial and temporal dynamics and hardware limitations. In this work, we introduce MaskFlow, a unified video generation framework that combines discrete representations with flow-matching to enable efficient generation of high-quality long videos. By leveraging a frame-level masking strategy during training, MaskFlow conditions on previously generated unmasked frames to generate videos with lengths ten times beyond that of the training sequences. MaskFlow does so very efficiently by enabling the use of fast Masked Generative Model (MGM)-style sampling and can be deployed in both fully autoregressive as well as full-sequence generation modes. We validate the quality of our method on the FaceForensics (FFS) and Deepmind Lab (DMLab) datasets and report Frechet Video Distance (FVD) competitive with state-of-the-art approaches. We also provide a detailed analysis on the sampling efficiency of our method and demonstrate that MaskFlow can be applied to both timestep-dependent and timestep-independent models in a training-free manner.

视频生成离散流长视频高效采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。