用完整帧替代掩码帧训练,提升视频生成连贯性
Taming Teacher Forcing for Masked Autoregressive Video Generation
- 用完整观察帧作为教师信号,取代传统掩码帧
- 首帧条件视频预测FVD提升23%,生成超100帧长视频
- 适合追求高质量长视频生成的研究者与开发者
我们提出MAGI,一种结合帧内掩码建模与帧间因果建模的混合视频生成框架。核心创新是完全教师强制(CTF),将掩码帧条件于完整观测帧而非掩码帧(即传统掩码教师强制MTF),实现从像素级到帧级自回归生成的平滑过渡。CTF显著优于MTF,在首帧条件视频预测上提升23%的FVD得分。为缓解暴露偏差,采用针对性训练策略,建立自回归视频生成新基准。实验表明,MAGI即使在仅16帧数据上训练,也能生成超过100帧的长且连贯视频序列,展现其在可扩展、高质量视频生成中的潜力。
原文摘要 · Abstract (English)
We introduce MAGI, a hybrid video generation framework that combines masked modeling for intra-frame generation with causal modeling for next-frame generation. Our key innovation, Complete Teacher Forcing (CTF), conditions masked frames on complete observation frames rather than masked ones (namely Masked Teacher Forcing, MTF), enabling a smooth transition from token-level (patch-level) to frame-level autoregressive generation. CTF significantly outperforms MTF, achieving a +23% improvement in FVD scores on first-frame conditioned video prediction. To address issues like exposure bias, we employ targeted training strategies, setting a new benchmark in autoregressive video generation. Experiments show that MAGI can generate long, coherent video sequences exceeding 100 frames, even when trained on as few as 16 frames, highlighting its potential for scalable, high-quality video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。