扩散模型在低维多模分布上统计最优,样本效率远超传统方法。
Diffusion Models Are Statistically Optimal for Learning Low-Dimensional Multi-Modal Distributions

- 基于数据内在低维结构设计理论分析框架
- 仅需约 $\widetilde{O}(\varepsilon^{-k \vee 2})$ 样本达 $\varepsilon$ 精度
- 无需平滑、有界密度等强假设,适合复杂多模数据
基于得分的扩散模型在学习高维分布方面表现出显著的实证成功,尤其适用于具有低维和多模结构的数据。然而,其统计效率的理论理解仍不充分。现有理论通常依赖于均匀有界密度或全局光滑得分函数等强正则性假设,难以捕捉这些内在结构。本文研究扩散模型在支持集为若干低维子空间并集时的样本复杂度。假设每个子空间内的数据分布为次高斯分布,我们证明扩散模型只需 $\widetilde{O}(\varepsilon^{-k \vee 2})$ 个样本即可在 1-沃瑟斯坦距离下达到 $\varepsilon$ 的误差,其中 $k$ 为内在维度。该接近最优的收敛速率仅依赖于内在维度,显著优于受维度灾难困扰的先前理论保证。值得注意的是,我们的分析适用于广泛分布类,无需平滑性、有界密度或对数凹性假设。总体而言,结果表明扩散模型可自适应地利用内在低维结构,同时自然处理多模数据,为其实现复杂高维学习任务提供了严格的理论支撑。
原文摘要 · Abstract (English)
Score-based diffusion models have demonstrated remarkable empirical success in learning high-dimensional distributions, particularly those exhibiting low-dimensional and multi-modal structures. However, theoretical understanding of their statistical efficiency remains limited. Existing theories typically rely on strong regularity assumptions, such as uniformly bounded densities or globally smooth score functions, which fail to capture such intrinsic structures. In this work, we study the sample complexity of diffusion models for learning distributions supported on a union of low-dimensional subspaces. Assuming that the data distribution within each subspace is subgaussian, we show that diffusion models require at most $\widetilde{O}(\varepsilon^{-k \vee 2})$ samples to achieve $\varepsilon$ error in 1-Wasserstein distance, where $k$ is the intrinsic dimension. This near-optimal convergence rate depends only on the intrinsic dimension and significantly improves upon prior theoretical guarantees that suffer from the curse of dimensionality. Notably, our analysis applies to a broad collection of distributions without imposing smoothness, bounded-density, or log-concavity assumptions. Overall, our results show that diffusion models can statistically adapt to intrinsic low-dimensional structure while naturally accommodating multi-modal data, offering a rigorous theoretical justification for their success in complex high-dimensional learning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。