arXiv:2606.02237cs.LG2026-06

发现扩散模型蒸馏中学生模型会自发复制教师的噪声-数据配对。

Why Are DMD Students Lazy? Understanding the Copying Behavior in Few-Step Distillation

论文配图:Why Are DMD Students Lazy? Understanding the Copying Behavior in Few-Step Distillation
图 1 · 摘自论文原文
  • 通过分布匹配实现多尺度蒸馏,学生可自由重映射潜在噪声。
  • 高维场景下学生自发复制教师原始噪声-数据配对,非偶然行为。
  • 该现象源于高维蒸馏中学生模型几何自由度受限,适合研究模型压缩者阅读。

分布匹配蒸馏(DMD)通过在所有尺度上对齐预训练扩散模型与学生模型的加噪分布,将大模型压缩为高效少步生成器。理论上,这种分布级监督不依赖教师模型的具体噪声-数据配对,使学生具备重映射潜在噪声的自由,在低维场景中已观察到此行为。然而,我们发现高维场景下,学生模型会自发复制教师的原始噪声-数据配对,这一现象称为复制。实验表明,复制并非对抗目标或教师记忆的副产物。我们的证据显示,复制是高维蒸馏过程中学生模型几何自由度受限所引发的涌现特性。

原文摘要 · Abstract (English)

Distribution Matching Distillation (DMD) compresses pretrained diffusion models into efficient few-step generators by aligning their noised distributions across all scales. In principle, such distribution-level supervision remains agnostic to specific noise-data pairings of the teacher; this provides the student the freedom to remap latent noise, a behavior consistently observed in low-dimensional settings. Surprisingly, we find that in high-dimensional settings, distilled students spontaneously reproduce the original noise-data pairings of the teacher, a phenomenon we term copying. We demonstrate that copying is neither a byproduct of adversarial objectives nor a result of teacher memorization. Instead, our evidence suggests that copying is an emergent property arising from the limited geometric freedom of the student model during high-dimensional distillation.

模型压缩扩散模型蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。