arXiv:2502.02922cs.LGcs.CV2025-02ICLR被引 2

提出可解析优化的预处理方法,显著加速一致性扩散模型训练

Elucidating the Preconditioning in Consistency Distillation

  • 基于教师ODE轨迹解析优化预处理系数
  • 实现2到3倍训练加速,提升学生轨迹对齐度
  • 适合研究扩散模型加速与一致性学习的学者

一致性蒸馏是加速扩散模型在一致性(轨迹)模型中应用的常用方法,其中学生模型被训练沿教师模型确定的概率流(PF)常微分方程(ODE)轨迹反向演化。预处理是稳定一致性蒸馏的关键技术,通过线性组合输入数据与网络输出,以预定义系数构成一致性函数,从而施加一致性边界条件,而不限制神经网络的形式与表达能力。然而,以往的预处理方案为手工设计,可能并非最优。本文首次从理论上揭示了预处理在一致性蒸馏中的设计准则及其与教师ODE轨迹的关系。基于此分析,我们提出一种名为 extit{Analytic-Precond}的系统化方法,可针对广义教师ODE上的一致性差距(定义为教师去噪器与最优学生去噪器之间的差距)进行解析优化。实验表明,Analytic-Precond能有效促进轨迹跳跃器的学习,增强学生轨迹与教师轨迹的一致性,在多种数据集上实现多步生成中2至3倍的训练加速。

原文摘要 · Abstract (English)

Consistency distillation is a prevalent way for accelerating diffusion models adopted in consistency (trajectory) models, in which a student model is trained to traverse backward on the probability flow (PF) ordinary differential equation (ODE) trajectory determined by the teacher model. Preconditioning is a vital technique for stabilizing consistency distillation, by linear combining the input data and the network output with pre-defined coefficients as the consistency function. It imposes the boundary condition of consistency functions without restricting the form and expressiveness of the neural network. However, previous preconditionings are hand-crafted and may be suboptimal choices. In this work, we offer the first theoretical insights into the preconditioning in consistency distillation, by elucidating its design criteria and the connection to the teacher ODE trajectory. Based on these analyses, we further propose a principled way dubbed \textit{Analytic-Precond} to analytically optimize the preconditioning according to the consistency gap (defined as the gap between the teacher denoiser and the optimal student denoiser) on a generalized teacher ODE. We demonstrate that Analytic-Precond can facilitate the learning of trajectory jumpers, enhance the alignment of the student trajectory with the teacher's, and achieve $2\times$ to $3\times$ training acceleration of consistency trajectory models in multi-step generation across various datasets.

扩散模型一致性蒸馏训练加速ODE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。