扩散模型无需精确拟合数据分布即可生成新样本,关键在于捕捉数据流形的几何结构。
Manifold Generalization Provably Proceeds Memorization in Diffusion Models
- 利用数据流形的几何特性,而非精确估计数据分布
- 生成能力的统计速率快于完整分布估计所需速率
- 特别适用于数据密度不规则的场景,适合生成模型研究者
扩散模型在仅学习粗粒度得分函数时仍能生成新样本,这一现象无法用传统密度估计视角解释。本文在流形假设下表明,粗粒度得分可捕捉数据支持集的几何结构,而忽略群体测度μ_data的细尺度分布。经典理论中,在k维流形上估计μ_data需达到 ilde{ m O}(N^{-1/k})的极小极大率;而使用粗粒度得分的扩散模型可利用流形的正则性,实现接近参数化的速率,逼近一个与μ_data在 ilde{ m O}(N^{-β/(4k)})邻域内密度相当的替代分布,其中β为流形正则性。该理论仅依赖底层支持的光滑性,尤其在数据密度不光滑(如不可导)时表现更优。当流形足够光滑时,生成新样本的泛化能力统计速率严格快于估计μ_data所需速率。
原文摘要 · Abstract (English)
Diffusion models often generate novel samples even when the learned score is only \emph{coarse} -- a phenomenon not accounted for by the standard view of diffusion training as density estimation. In this paper, we show that, under the \emph{manifold hypothesis}, this behavior can instead be explained by coarse scores capturing the \emph{geometry} of the data while discarding the fine-scale distributional structure of the population measure~$μ_{\scriptscriptstyle\mathrm{data}}$. Concretely, whereas estimating the full data distribution $μ_{\scriptscriptstyle\mathrm{data}}$ supported on a $k$-dimensional manifold is known to require the classical minimax rate $\tilde{\mathcal{O}}(N^{-1/k})$, we prove that diffusion models trained with coarse scores can exploit the \emph{regularity of the manifold support} and attain a near-parametric rate toward a \emph{different} target distribution. This target distribution has density uniformly comparable to that of~$μ_{\scriptscriptstyle\mathrm{data}}$ throughout any $\tilde{\mathcal{O}}\bigl(N^{-β/(4k)}\bigr)$-neighborhood of the manifold, where $β$ denotes the manifold regularity. Our guarantees therefore depend only on the smoothness of the underlying support, and are especially favorable when the data density itself is irregular, for instance non-differentiable. In particular, when the manifold is sufficiently smooth, we obtain that \emph{generalization} -- formalized as the ability to generate novel, high-fidelity samples -- occurs at a statistical rate strictly faster than that required to estimate the full population distribution~$μ_{\scriptscriptstyle\mathrm{data}}$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。