通过分尺度蒸馏,让扩散模型用极少步数生成高质量图像视频。
Scale-wise Distillation of Diffusion Models
- 按尺度分阶段蒸馏,避免中间步骤重复计算。
- 仅需约2步全分辨率采样,质量超越现有方法。
- 适合追求高效生成的图像/视频生成研究者。
近期扩散模型蒸馏方法已实现显著进展,使大规模文本条件图像与视频扩散模型在约4步内完成高质量采样。然而,进一步减少采样步数愈发困难,提示效率提升应转向其他模型维度。为此,我们提出SwD——一种分尺度扩散蒸馏框架,使少步数模型具备渐进式生成能力,避免在中间扩散时间步进行冗余计算。除提升效率外,SwD通过基于最大均值差异(MMD)的简单块级蒸馏目标,丰富了分布匹配蒸馏方法族。该目标显著改善现有蒸馏方法的收敛性,在孤立使用时表现也出人意料地优秀,可作为扩散蒸馏的有力基线。应用于当前最先进的文本到图像/视频扩散模型,SwD在相同计算预算下逼近2步全分辨率采样的速度,并在自动指标与人工偏好评估中显著优于现有方案。
原文摘要 · Abstract (English)
Recent diffusion distillation methods have achieved remarkable progress, enabling high-quality ${\sim}4$-step sampling for large-scale text-conditional image and video diffusion models. However, further reducing the number of sampling steps becomes more and more challenging, suggesting that efficiency gains may be better mined along other model axes. Motivated by this perspective, we introduce SwD, a scale-wise diffusion distillation framework that equips few-step models with progressive generation, avoiding redundant computations at intermediate diffusion timesteps. Beyond efficiency, SwD enriches the family of distribution matching distillation approaches by introducing a simple patch-level distillation objective based on Maximum Mean Discrepancy (MMD). This objective significantly improves the convergence of existing distillation methods and performs surprisingly well in isolation, offering a competitive baseline for diffusion distillation. Applied to state-of-the-art text-to-image/video diffusion models, SwD approaches the sampling speed of two full-resolution steps and largely outperforms alternatives under the same compute budget, as evidenced by automatic metrics and human preference studies. Project page: https://yandex-research.github.io/swd
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。