arXiv:2605.20235cs.LGcs.AI2026-05

提出新方法让扩散模型高效学习低维流形数据,理论证明可突破维度诅咒。

Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine

论文配图:Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine
图 1 · 摘自论文原文
  • 基于得分函数几何设计分阶段的降维与密度优化机制
  • 在多个数据集上生成质量优于或等同于传统方法,重建更准确
  • 适合研究扩散模型理论、流形学习或分子生成的学者

扩散模型能高质量生成高维数据,但当数据位于低维流形时,其训练如何有效学习得分函数并避开维度诅咒仍缺乏理论解释。我们发现得分函数本身的几何特性驱动了‘坍缩-精炼’机制:在小噪声尺度下,得分函数的奇异发散使去噪映射快速坍缩至数据流形投影;在中等噪声尺度下,训练精炼流形上的内在密度分布。我们据此构建了得分诱导潜空间扩散(SiLD),一种两阶段框架,仅通过单一去噪得分匹配目标同时实现流形学习与密度估计,取代了基于VAE的潜空间扩散模型中的启发式KL正则化。理论证明,样本复杂度依赖于内在维度而非环境维度。在堆叠MNIST、CelebA变体和分子生成基准上的实验表明,SiLD在生成质量上达到或超越基于VAE的LDMs,且重建性能始终更优,验证了理论预测。

原文摘要 · Abstract (English)

Diffusion models generate high-dimensional data with remarkable quality, yet how their training efficiently learns the score function, bypassing the curse of dimensionality when data is supported on low-dimensional manifolds, remains theoretically unexplained. We identify a collapse-and-refine mechanism driven by the geometry of the score function itself: at small noise scales, the diverging singularity of the score drives a rapid dimensional collapse of the induced denoising map onto the data manifold projection; at moderate noise scales, training refines the intrinsic density on the learned manifold. We instantiate this principle as Score-induced Latent Diffusion (SiLD), a two-stage framework in which both manifold learning and density estimation emerge from a single denoising score matching objective, replacing the heuristic KL regularization of VAE-based latent diffusion models. We prove that the resulting sample complexity depends on the intrinsic dimension rather than the ambient dimension. Experiments on Stacked MNIST, CelebA variants, and molecular generation benchmarks show that SiLD matches or outperforms VAE-based LDMs in generation quality and consistently improves reconstruction, validating our theoretical predictions.

扩散模型流形学习得分匹配生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。