通过噪声级别专家分工,实现两步视频生成的质量与多样性兼得。
DUET: A Diversity-Quality Duet of Distillation Experts for Two-Step Video Generation

- 高噪声步用sCM专家布局结构,低噪声步用DMD专家细化细节。
- 两步生成质量接近DMD,结构多样性约为DMD的两倍。
- 适合追求高效高质量视频生成的研究者和开发者。
近年来,扩散模型实现了高质量视频生成,但迭代采样成本过高,制约了实际应用。少步蒸馏虽缓解了成本问题,却暴露了质量与多样性的权衡:轨迹级蒸馏(如sCM)侧重多样性,分布级蒸馏(如DMD)侧重质量。针对极简两步视频生成,我们提出DUET,通过噪声级别专家协作:高噪声步由sCM专家负责布局多样性结构,低噪声步由DMD专家负责精细化外观。两者独立训练,避免损失组合优化难题,实现质量与多样性的联合提升。进一步识别出中继接口与高噪声阶段为瓶颈,引入强化学习引导的专家自适应,形成DUET+。基于Wan2.1-T2V-1.3B骨干网络,DUET使sCM的两步生成质量接近DMD水平,同时保留近两倍于DMD的结构多样性;DUET+进一步提升整体质量并维持该多样性优势。结果表明,噪声级别专家专业化是两步视频生成中兼顾质量与多样性的有效范式。
原文摘要 · Abstract (English)
Diffusion models have enabled high-quality video generation in recent years, but the high cost of iterative sampling hinders their practical deployment. Few-step distillation alleviates this cost, yet exposes a quality--diversity trade-off between its two dominant paradigms: trajectory-level distillation (e.g., sCM) favors diversity, whereas distribution-level distillation (e.g., DMD) favors quality. Targeting extreme two-step video generation, we introduce DUET, which reconciles the two paradigms through a noise-level duet of experts: an sCM expert takes the high-noise step to lay out diverse structure, and a DMD expert takes the low-noise step to refine appearance detail. Since the two experts are trained independently with their native objectives, DUET sidesteps the optimization difficulties of loss-level combinations and delivers quality and diversity jointly rather than trading one for the other. We further identify the relay interface and the high-noise stage as the remaining bottlenecks, and address them with RL-guided expert adaptation, yielding DUET+. With the Wan2.1-T2V-1.3B backbone, DUET lifts the two-step quality of sCM close to the level of DMD while retaining nearly all of its structural diversity---about twice that of DMD---and DUET+ further improves overall quality while preserving this diversity advantage. Together, these results establish noise-level expert specialization as a simple, effective paradigm for reconciling diversity and quality in two-step video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。