解决快速图像生成中样本多样性下降问题,无需额外模块或正则化。
Diversity-Preserved Distribution Matching Distillation for Fast Visual Synthesis
- 分阶段蒸馏:前期用目标预测保多样性,后期用标准损失提画质。
- 仅需4步采样即保持高多样性与竞争力画质,超越现有方法。
- 无对抗训练、无额外模块,训练稳定且易扩展,适合实际部署。
分布匹配蒸馏(DMD)通过将学生模型对齐参考多步教师模型,实现少步图像生成。然而,实践中优化DMD会降低少步合成的样本多样性,现有修复方法通常依赖感知或对抗正则化,导致训练不稳定且难以扩展。本文提出多样性保留的DMD(DP-DMD),受早期与晚期去噪步骤互补作用启发,采用角色分离策略:第一阶段使用教师引导的目标预测目标(如v-prediction)以保持样本多样性,其余阶段则使用标准DMD损失优化感知质量。DP-DMD无需感知或对抗正则化,无需额外模块,也无需教师生成的参考样本,在少步采样下仍能保持高多样性并达到有竞争力的视觉质量,提供一种简单且稳定的DMD替代方案。
原文摘要 · Abstract (English)
Distribution matching distillation (DMD) facilitates few-step image generation by aligning a distilled student with a reference multi-step teacher. In practice, however, optimizing DMD can reduce sample diversity in few-step synthesis, and existing remedies typically rely on perceptual or adversarial regularization, leading to stability and scalability challenges during training. Here, we describe diversity-preserved DMD (DP-DMD), a role-separated distillation method inspired by the complementary roles of early and late denoising steps. Specifically, the first distillation step is trained with a teacher-derived target-prediction objective (e.g., v-prediction) to preserve sample diversity, while the remaining steps are optimized with the standard DMD loss to refine perceptual quality. DP-DMD, with no perceptual or adversarial regularization, no additional modules, and no teacher-generated reference samples, preserves sample diversity while maintaining competitive visual quality under few-step sampling, providing a simple and stable alternative to other DMD variants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。