arXiv:2505.20123cs.LGcs.CV2025-05被引 1

提出概率流距离度量扩散蒸馏泛化能力,揭示训练规律。

Understanding Generalization in Diffusion Distillation via Probability Flow Distance

  • 用概率流ODE映射差异定义新度量PFD,理论严谨且计算高效。
  • 发现从记忆到泛化的量化演化规律,呈现双下降训练动态。
  • 为扩散模型泛化研究提供可计算框架,适合模型评估与优化者。

扩散蒸馏能有效训练轻量级、少步数的扩散模型以实现高效生成。然而,其泛化性能评估仍具挑战:理论指标对高维数据不实用,而缺乏实际可行的严格度量。本文通过引入概率流距离(PFD),建立了一种理论基础扎实且计算高效的泛化度量方法。PFD通过比较由概率流微分方程(ODE)诱导的噪声到数据映射来量化分布间距离。在扩散蒸馏设定下,利用PFD实证揭示了若干关键泛化行为:(1)从记忆到泛化的定量演化规律;(2)每轮训练中的双下降动态;(3)偏差-方差分解。本工作为扩散蒸馏的泛化研究奠定基础,并将其与扩散训练研究相连接。

原文摘要 · Abstract (English)

Diffusion distillation provides an effective approach for learning lightweight and few-steps diffusion models with efficient generation. However, evaluating their generalization remains challenging: theoretical metrics are often impractical for high-dimensional data, while no practical metrics rigorously measure generalization. In this work, we bridge this gap by introducing probability flow distance (\texttt{PFD}), a theoretically grounded and computationally efficient metric to measure generalization. Specifically, \texttt{PFD} quantifies the distance between distributions by comparing their noise-to-data mappings induced by the probability flow ODE. Using \texttt{PFD} under the diffusion distillation setting, we empirically uncover several key generalization behaviors, including: (1) quantitative scaling behavior from memorization to generalization, (2) epoch-wise double descent training dynamics, and (3) bias-variance decomposition. Beyond these insights, our work lays a foundation for generalization studies in diffusion distillation and bridges them with diffusion training.

扩散模型泛化分析概率流蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。