arXiv:2605.07969cs.LGcs.IT2026-05

揭示扩散模型为何能在高维数据中高效采样。

When Diffusion Model Can Ignore Dimension: An Entropy-Based Theory

  • 基于信息论,用潜变量熵替代维度控制采样误差。
  • 高维数据采样复杂度仅与潜熵线性相关,与维度无关。
  • 适用于自然图像等具有紧凑潜表示的数据,适合研究生成模型理论者。

扩散模型在高维数据(如图像)上表现优异,通常仅需少量反向时间步即可生成高质量样本。尽管如此,现有收敛理论未能充分解释其在高维下的效率。多数已有KL界依赖于环境维度,而改进结果则引入内在维度或几何结构假设。本文提出一种新的信息论视角:对于高斯混合目标分布,离散化误差由潜在混合分量的香农熵决定,而非环境维度。因此,主要步数复杂度与潜熵线性相关,且对数据的二阶矩仅呈对数依赖。该分析还扩展至离散目标分布,此时复杂度由目标熵决定,而非嵌入空间维度。结果表明,当数据分布具备紧凑潜表示时(如自然图像),扩散采样可在高维空间保持高效。

原文摘要 · Abstract (English)

Diffusion models perform remarkably well on high-dimensional data such as images, often using only a modest number of reverse-time steps. Despite this practical success, existing convergence theory does not fully explain why such samplers remain efficient in high dimensions. Many prior KL guarantees bound the discretization error in terms of the ambient dimension, while other improved results replace this dependence using intrinsic-dimensional or geometric structure assumptions. In this work, we develop an alternative information-theoretic perspective on diffusion sampler convergence. We prove that, for Gaussian mixture targets, the discretization error is controlled by the Shannon entropy of the latent mixture component rather than by the ambient dimension. Consequently, the leading step complexity scales linearly with latent entropy and depends only logarithmically on the second moment of the data. Our analysis also extends to discrete target distributions, where the relevant complexity is the entropy of the target rather than the dimension of the embedding space. These results suggest that diffusion sampling can remain efficient in high-dimensional spaces when the data distribution admits a compact latent representation, as is widely believed to be the case for natural images.

扩散模型信息论采样效率潜变量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。