arXiv:2603.03700stat.MLcs.AI2026-03被引 6

扩散模型能自动适应数据内在维度,突破高维瓶颈。

Generalization Properties of Score-matching Diffusion Models for Intrinsically Low-dimensional Data

  • 基于分数匹配的扩散模型学习分布,理论证明收敛速度与数据内在维度相关。
  • 在仅需有限阶矩的条件下,误差率随样本数呈幂律下降,依赖于$(p,q)$-Wasserstein维数。
  • 适用于图像等低维流形数据,为扩散模型提供可解释的统计保障,适合研究生成模型理论者。

尽管基于分数的扩散模型在实践中表现优异,但其统计保证仍不充分。现有分析常给出悲观的收敛速率,未能反映真实数据中普遍存在的内在低维结构(如自然图像)。本文研究从有限样本中学习未知分布 $μ$ 的分数匹配扩散模型的统计收敛性。在对前向扩散过程和数据分布施加弱正则性条件的前提下,我们推导出学习到的生成分布与真实分布之间在 Wasserstein-$p$ 距离下的有限样本误差界。不同于以往结果,我们的保证适用于所有 $p \ge 1$,且仅要求 $μ$ 具有有限 $q$-阶矩,无需紧支撑、流形或光滑密度假设。具体而言,在给定 $n$ 个来自 $μ$ 的独立同分布样本、恰当选择网络架构、超参数与离散化方案下,我们证明期望 Wasserstein-$p$ 误差满足 $\mathbb{E}\, \mathbb{W}_p(\hatμ,μ) = \widetilde{O}\left(n^{-1 / d^\ast_{p,q}(μ)}\right)$,其中 $d^\ast_{p,q}(μ)$ 是 $μ$ 的 $(p,q)$-Wasserstein 维数。结果表明,扩散模型能自然适应数据内在几何结构,缓解维度灾难,因为收敛速率取决于 $d^\ast_{p,q}(μ)$ 而非环境维度。此外,该理论概念上连接了扩散模型与 GAN 的分析,以及最优传输中建立的尖锐极小极大率。提出的 $(p,q)$-Wasserstein 维数还扩展了经典 Wasserstein 维数至无界支撑分布,具有独立理论价值。

原文摘要 · Abstract (English)

Despite the remarkable empirical success of score-based diffusion models, their statistical guarantees remain underdeveloped. Existing analyses often provide pessimistic convergence rates that do not reflect the intrinsic low-dimensional structure common in real data, such as that arising in natural images. In this work, we study the statistical convergence of score-based diffusion models for learning an unknown distribution $μ$ from finitely many samples. Under mild regularity conditions on the forward diffusion process and the data distribution, we derive finite-sample error bounds on the learned generative distribution, measured in the Wasserstein-$p$ distance. Unlike prior results, our guarantees hold for all $p \ge 1$ and require only a finite-moment assumption on $μ$, without compact-support, manifold, or smooth-density conditions. Specifically, given $n$ i.i.d.\ samples from $μ$ with finite $q$-th moment and appropriately chosen network architectures, hyperparameters, and discretization schemes, we show that the expected Wasserstein-$p$ error between the learned distribution $\hatμ$ and $μ$ scales as $\mathbb{E}\, \mathbb{W}_p(\hatμ,μ) = \widetilde{O}\!\left(n^{-1 / d^\ast_{p,q}(μ)}\right),$ where $d^\ast_{p,q}(μ)$ is the $(p,q)$-Wasserstein dimension of $μ$. Our results demonstrate that diffusion models naturally adapt to the intrinsic geometry of data and mitigate the curse of dimensionality, since the convergence rate depends on $d^\ast_{p,q}(μ)$ rather than the ambient dimension. Moreover, our theory conceptually bridges the analysis of diffusion models with that of GANs and the sharp minimax rates established in optimal transport. The proposed $(p,q)$-Wasserstein dimension also extends the notion of classical Wasserstein dimension to distributions with unbounded support, which may be of independent theoretical interest.

扩散模型统计分析生成模型低维结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。