通过简单添加惯性更新,解决高维流形数据扩散模型的过拟合问题。
Resolving Memorization in Empirical Diffusion Model for Manifold Data in High-Dimensional Spaces
- 在扩散过程末尾引入惯性更新,无需额外训练。
- 生成样本与真实分布的Wasserstein-1距离为O(n^{-2/(d+4)})。
- 适用于高维流形数据生成,特别适合追求多样性新样本的研究者。
扩散模型通过加噪和去噪过程生成新数据,但在数据分布由n个点构成时,常会重复已有样本,即记忆效应。现有方法多依赖复杂机器学习技术,本文提出仅需在经验扩散模拟末尾加入惯性更新即可解决该问题。所提惯性扩散模型仅需经验得分函数,无需额外训练。我们证明其生成样本在C²流形上逼近真实分布,且与总体分布的Wasserstein-1距离为O(n^{-2/(d+4)})。该界显著缩小了分布间距离,验证了模型能生成新颖多样样本。值得注意的是,此估计与环境空间维度无关,因无需进一步训练。分析揭示惯性扩散样本类似流形上的高斯核密度估计,建立了扩散模型与流形学习的新关联。
原文摘要 · Abstract (English)
Diffusion models are popular tools for generating new data samples, using a forward process that adds noise to data and a reverse process to denoise and produce samples. However, when the data distribution consists of n points, empirical diffusion models tend to reproduce existing data points, a phenomenon known as the memorization effect. Current literature often addresses this with complex machine learning techniques. This work shows that the memorization issue can be solved simply by applying an inertia update at the end of the empirical diffusion simulation. Our inertial diffusion model requires only the empirical score function and no additional training. We demonstrate that the distribution of samples from this model approximates the true data distribution on a $C^2$ manifold of dimension $d$, within a Wasserstein-1 distance of order $O(n^{-\frac{2}{d+4}})$. This bound significantly shrinks the Wasserstein distance between the population and empirical distributions, confirming that the inertial diffusion model produces new and diverse samples. Remarkably, this estimate is independent of the ambient space dimension, as no further training is needed. Our analysis shows that the inertial diffusion samples resemble Gaussian kernel density estimations on the manifold, revealing a novel connection between diffusion models and manifold learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。