通过双噪声掩码机制,提升自回归视频生成的清晰度与真实感。
Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout

- 在自回归扩散蒸馏中引入时空随机掩码,扰动生成过程以避免模式坍塌。
- 显著改善视频视觉质量,解决过饱和与过度平滑问题,无需额外数据或训练阶段。
- 适用于追求高真实感视频生成的研究者与开发者,尤其适合轻量化部署场景。
自回归(AR)视频扩散模型在实时视频生成方面展现出巨大潜力。现有方法通过分布匹配蒸馏(DMD)将预训练的双向视频扩散模型压缩为因果AR学生模型,但生成视频常出现过饱和和过度平滑问题,导致视觉质量与真实感不足。其关键原因在于DMD中反向KL目标引发的模式搜索行为,使学生分布坍缩至教师分布的少数模式。为此,本文提出掩码强制(Mask Forcing),一种双噪声掩码滚动策略,通过在自回归学生模型的滚动生成过程中沿时空轴注入随机掩码,对噪声输入进行清洁信号扰动。该机制促使学生生成探索教师分布更多区域,使DMD提供超出学生已有覆盖模式的学习信号。同时,更清晰的令牌为其他噪声令牌提供去噪引导,改善中间生成预测,减少误差累积。大量实验表明,该方法可有效提升多种AR视频扩散蒸馏方法的视觉质量,且不依赖真实视频数据或额外后训练阶段。
原文摘要 · Abstract (English)
Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Matching Distillation (DMD), but the generated videos often suffer from over-saturation and over-smoothing issues, resulting in limited visual quality and realism. The key contributing factor is the mode-seeking behavior of the reverse KL objective in DMD, which can cause the student distribution to collapse onto only a few modes of the teacher distribution. To address this, we propose Mask Forcing, a Dual-Noise Masking Rollout strategy that perturbs the AR student self-rollout to mitigate mode collapse induced by reverse-KL mode seeking. The core idea is to inject cleaner signals into noisy rollout inputs via random masks along spatial and temporal axes during the self-rollout process of AR diffusion distillation. Such perturbations encourage the student rollouts to explore more regions of the teacher distribution, allowing DMD to provide learning signals beyond the modes already covered by the student. Moreover, the cleaner tokens act as denoising guidance for other noisier tokens, improving the intermediate rollout predictions and reducing error accumulation. Extensive experiments demonstrate that our method improves multiple AR video diffusion distillation methods with higher visual quality efficiently, without incorporating real video data or additional post-training stages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。