提出扩散模型生成分布与真实数据间Wasserstein距离的紧界,揭示其采样复杂度随维度增长缓慢。
Wasserstein Bounds for generative diffusion models with Gaussian tail targets
- 基于数据尾部服从高斯型假设,推导得分模型全局Lipschitz性质
- 证明采样复杂度为O(√d),对数常数项小,与协方差迹线性相关
- 适用于早期停止等实际场景,理论支撑扩散模型高效采样
我们给出了得分驱动生成模型生成分布与真实数据分布之间Wasserstein距离的估计。在数据分布具有高斯型尾部行为且得分函数ε-准确近似的情况下,采样复杂度关于维度为$\mathcal{O}(\sqrt{d})$,包含一个对数常数。该高斯尾假设具有普遍性,可涵盖早期停止技术中具有有界支持的实际目标分布。分析核心在于通过热核的维数无关估计,获得得分函数的全局Lipschitz界。因此,我们的复杂度上界与协方差算子迹的平方根呈线性关系(最多含对数因子),这与前向过程的不变分布相关。
原文摘要 · Abstract (English)
We present an estimate of the Wasserstein distance between the data distribution and the generation of score-based generative models. The sampling complexity with respect to dimension is $\mathcal{O}(\sqrt{d})$, with a logarithmic constant. In the analysis, we assume a Gaussian-type tail behavior of the data distribution and an $ε$-accurate approximation of the score. Such a Gaussian tail assumption is general, as it accommodates a practical target - the distribution from early stopping techniques with bounded support. The crux of the analysis lies in the global Lipschitz bound of the score, which is shown from the Gaussian tail assumption by a dimension-independent estimate of the heat kernel. Consequently, our complexity bound scales linearly (up to a logarithmic constant) with the square root of the trace of the covariance operator, which relates to the invariant distribution of the forward process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。