arXiv:2505.16024cs.LGcs.AI2025-05被引 3

揭示扩散模型轨迹蒸馏的理论机制,指导高效采样方法设计。

Toward Theoretical Insights into Diffusion Trajectory Distillation via Operator Merging

  • 将轨迹蒸馏重新理解为算子合并问题,从理论层面分析其原理。
  • 线性高斯场景下信号衰减是主要瓶颈,可推导出最优合并策略。
  • 非线性混合场景中合并会引入不可避误差,且误差随步骤指数增长。

扩散轨迹蒸馏通过训练学生模型以较少步数逼近预训练教师模型的多步去噪轨迹,从而加速采样。尽管实证效果显著,但蒸馏策略与生成质量之间的权衡仍不清晰。本文通过将轨迹蒸馏重释为算子合并问题,区分两种不同情形进行理论分析。在线性高斯情形下,近似误差为零,优化误差(尤其是由有限训练时间引起的信号衰减)成为主要瓶颈,由此推导出理论上最优的合并策略,该策略具有方差驱动的相变特性,并可通过帕累托动态规划算法计算。在非线性高斯混合情形下,我们证明合并复合步骤会引发不可避免的近似误差,且因混合成分呈指数增长而加剧误差传播。上述结果澄清了两类情形下的理论机制,为方法选择提供了原则性指导。

原文摘要 · Abstract (English)

Diffusion trajectory distillation accelerates sampling by training a student model to approximate the multi-step denoising trajectories of a pretrained teacher model using far fewer steps. Despite strong empirical results, the trade-off between distillation strategy and generative quality remains poorly understood. We provide a theoretical characterization by reinterpreting trajectory distillation as an operator merging problem, differentiating our analysis between two distinct regimes. In the linear Gaussian regime, where approximation error is zero, we isolate optimization error, specifically signal shrinkage driven by finite training time, as the primary bottleneck. This characterization allows us to derive the theoretically optimal merging strategy, which exhibits a variance-driven phase transition and is computable via a Pareto dynamic programming algorithm. In the nonlinear Gaussian mixture regime, we prove that distilling composite steps incurs unavoidable approximation error due to the exponential growth of mixture components, and we quantify how these errors amplify across merges. Together, these results clarify the distinct theoretical mechanisms governing each regime and provide principled guidance for method selection.

扩散模型轨迹蒸馏理论分析算子合并

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。