arXiv:2602.15971cs.LGcs.AI2026-02

通过多分支对齐中间轨迹,提升扩散模型蒸馏效率与生成质量

B-DENSE: Branching For Dense Ensemble Network Supervision Efficiency

  • 设计多分支结构,让学生模型同时学习教师模型各中间步骤
  • 相比传统蒸馏,生成图像质量显著提升,减少离散化误差
  • 适合关注高效扩散模型训练与高质量图像生成的研究者

受非平衡热力学启发,扩散模型在生成建模中取得了顶尖性能。然而其迭代采样特性导致推理延迟高。尽管近期蒸馏技术可加速采样,但会丢弃中间轨迹步骤,造成结构信息损失并引入显著离散化误差。为此,我们提出 B-DENSE 框架,利用多分支轨迹对齐。将学生模型架构修改为输出 K 倍扩展通道,每个子集对应教师轨迹中一个离散中间步骤。通过训练这些分支同时映射到教师目标时间步的完整序列,强制实现密集中间轨迹对齐。由此,学生模型从训练初期即学会导航解空间,在图像生成质量上优于基线蒸馏框架。

原文摘要 · Abstract (English)

Inspired by non-equilibrium thermodynamics, diffusion models have achieved state-of-the-art performance in generative modeling. However, their iterative sampling nature results in high inference latency. While recent distillation techniques accelerate sampling, they discard intermediate trajectory steps. This sparse supervision leads to a loss of structural information and introduces significant discretization errors. To mitigate this, we propose B-DENSE, a novel framework that leverages multi-branch trajectory alignment. We modify the student architecture to output $K$-fold expanded channels, where each subset corresponds to a specific branch representing a discrete intermediate step in the teacher's trajectory. By training these branches to simultaneously map to the entire sequence of the teacher's target timesteps, we enforce dense intermediate trajectory alignment. Consequently, the student model learns to navigate the solution space from the earliest stages of training, demonstrating superior image generation quality compared to baseline distillation frameworks.

扩散模型模型蒸馏图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。