改进扩散模型生成过程的KL散度收敛分析,提升精度与效率。
A Sharp KL-Convergence Analysis for Diffusion Models under Minimal Assumptions
- 将生成过程拆分为反向ODE与微小加噪两步,优化误差控制。
- 在无光滑性假设下,实现线性于数据维度d的误差依赖。
- 适合关注生成模型理论分析的研究者,尤其对数学推导严谨性有要求者。
基于扩散的生成模型已成为生成高质量样本的有效方法。近期工作聚焦于在最小假设下分析其生成过程的收敛性,通过反向SDE或概率流ODE实现。目前最优的无光滑性假设下的KL散度保证为:与数据维度d呈线性关系,与ε呈反二次关系。本文提出更精细的分析,改进了对ε的依赖。我们将生成过程建模为两个步骤的组合:先进行反向ODE步骤,再沿前向过程执行一次较小的加噪步骤。该设计利用了反向ODE步骤在Wasserstein型误差上的可控性,通过添加噪声将其转化为KL散度上界,从而获得对离散化步长更优的依赖。此外,我们提出一种新分析方法,在无光滑性假设下实现概率流ODE离散化误差的线性d依赖。我们证明:当目标分布被方差为δ的高斯噪声污染时,仅需$ ilde{O}ig( frac{d\ ext{log}^{3/2}(1/δ)}{\varepsilon}ig)$步即可在KL散度中达到$O(\varepsilon^2)$的误差,优于先前最佳结果所需的$ ilde{O}ig( frac{d\text{log}^2(1/δ)}{\varepsilon^2}ig)$步。
原文摘要 · Abstract (English)
Diffusion-based generative models have emerged as highly effective methods for synthesizing high-quality samples. Recent works have focused on analyzing the convergence of their generation process with minimal assumptions, either through reverse SDEs or Probability Flow ODEs. The best known guarantees, without any smoothness assumptions, for the KL divergence so far achieve a linear dependence on the data dimension $d$ and an inverse quadratic dependence on $\varepsilon$. In this work, we present a refined analysis that improves the dependence on $\varepsilon$. We model the generation process as a composition of two steps: a reverse ODE step, followed by a smaller noising step along the forward process. This design leverages the fact that the ODE step enables control in Wasserstein-type error, which can then be converted into a KL divergence bound via noise addition, leading to a better dependence on the discretization step size. We further provide a novel analysis to achieve the linear $d$-dependence for the error due to discretizing this Probability Flow ODE in absence of any smoothness assumptions. We show that $\tilde{O}\left(\tfrac{d\log^{3/2}(\frac{1}δ)}{\varepsilon}\right)$ steps suffice to approximate the target distribution corrupted with Gaussian noise of variance $δ$ within $O(\varepsilon^2)$ in KL divergence, improving upon the previous best result, requiring $\tilde{O}\left(\tfrac{d\log^2(\frac{1}δ)}{\varepsilon^2}\right)$ steps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。