用简单方法修复视频生成模型多样性与真实感缺失问题
Data-Forcing Distillation: Restoring Diversity and Fidelity in Few-Step Video Generation

- 引入教师模型得分差异引导学生学习真实数据分布
- 仅需100-300步微调,显著提升生成视频的多样性和真实感
- 适用于文本到视频、图像到视频等多种生成任务
近期研究将多步视频扩散模型蒸馏为高效少步学生模型,其中分布匹配蒸馏(DMD)及其改进版DMD2实现了高质量生成与快速收敛。然而,由于反向KL目标的固有特性,这些方法存在两个持续性缺陷:样本多样性大幅下降,输出明显过饱和且偏离真实视频外观。本文提出数据强制蒸馏(DFD),一种仅需一行代码修改的后训练框架,可有效恢复DMD的多样性与保真度。其核心思想是利用教师模型得分差异,引导学生向真实数据分布靠拢,从而填补缺失模式(缓解模式崩溃),避开真实数据中不存在的问题模式(避免过饱和)。我们提供了理论分析,并在文本到视频、图像到视频及自回归视频生成任务上验证了该方法。仅需100–300步微调,DFD在Wan2.1-1.3B和Cosmos-Predict2.5-2B模型上均显著恢复了生成多样性与真实性,消除过饱和伪影,视频动态与外观更优,甚至超越教师模型。
原文摘要 · Abstract (English)
Recent progress has shown promise in distilling multi-step video diffusion models into efficient few-step students. Among them, Distribution Matching Distillation (DMD) and its successor DMD2 achieved strong generation quality and fast convergence. However, due to the nature of the reverse Kullback--Leibler (KL) objective, these methods exhibit two persistent failure modes: a substantial drop in sample diversity, and visibly over-saturated outputs that deviate from real-video appearance. In this work, we propose Data-Forcing Distillation (DFD), a simple post-training framework that restores diversity and fidelity in DMD with only a single-line of code change. At its core is the teacher score discrepancy to guide the student toward the real-data distribution, pulling it to missing modes (mitigating mode collapse) and away from problematic modes absent in real data (avoiding over-saturation). We provide an in-depth theoretical analysis of our framework and validate our approach on text-to-video, image-to-video, and autoregressive video generation. With only 100--300 steps of finetuning, DFD effectively restores diversity and fidelity on both Wan2.1-1.3B and Cosmos-Predict2.5-2B model, resolving the over-saturation artifacts with significantly better video dynamics and appearance, and even outperforms the teacher model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。