arXiv:2409.02426cs.LGcs.CV2024-09被引 61

扩散模型能高效学习低维数据分布,突破维度诅咒。

Breaking the Curse of Dimensionality: Diffusion Models Efficiently Learn Low-Dimensional Distributions

  • 将数据建模为低秩高斯混合,证明训练等价于子空间聚类
  • 样本复杂度与数据内在维度线性相关,而非环境维度指数增长
  • 学到的子空间对应图像语义属性,支持可控生成

尽管扩散模型在各类生成任务中表现优异,但其学习数据分布的基本原理仍不清晰。本文提出新数学框架,解释扩散模型如何从有限样本中有效学习低维分布,避免维度诅咒。基于图像数据的内在低维结构,我们理论分析了数据分布为低秩高斯混合的情形。在合适网络参数化下,扩散模型的训练目标等价于对训练样本进行标准子空间聚类,每个子空间基对应一个高斯成分的低秩协方差。该等价性表明,学习底层分布的样本复杂度与数据内在维度呈线性关系,而非环境维度的指数增长。理论结果得到合成及真实图像数据集上泛化相变现象的实证支持。此外,我们建立了学习到的子空间基与图像语义属性之间的对应关系,为可控图像生成提供了理论基础。

原文摘要 · Abstract (English)

Despite their empirical success across a wide range of generative tasks, the fundamental principles underlying the ability of diffusion models to learn data distributions are poorly understood. In this work, we develop a new mathematical framework that explains how diffusion models can effectively learn low-dimensional distributions from a finite number of training samples without suffering from the curse of dimensionality. Specifically, motivated by the intrinsic low-dimensional structure of image data, we theoretically analyze a setting in which the data distribution is modeled as a mixture of low-rank Gaussians. Under suitable network parameterization, we show that optimizing the training objective of diffusion models is equivalent to solving the canonical subspace clustering problem over the training samples, where each subspace basis corresponds to the low-rank covariance of a Gaussian component. This equivalence allows us to show that the sample complexity for learning the underlying distribution scales linearly with the intrinsic dimension of the data, rather than exponentially with the ambient dimension. Our theoretical findings are further supported by empirical evidence that demonstrates phase transition phenomena in generalization on both synthetic and real-world image datasets. Moreover, we establish a correspondence between the learned subspace bases and semantic attributes of image data, providing a principled foundation for controllable image generation.

扩散模型低维结构子空间聚类可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。