arXiv:2502.00336cs.LGstat.ML2025-02被引 17

解析扩散模型泛化与记忆机制,揭示训练样本数与噪声采样数的影响。

Denoising Score Matching with Random Features: Insights on Diffusion Models from Precise Learning Curves

  • 用随机特征网络建模得分函数,推导测试与训练误差的精确表达式。
  • 发现泛化与记忆由数据量比、特征数比和噪声采样数共同决定。
  • 理论结果匹配实验现象,为扩散模型设计提供定量指导。

我们从理论上研究扩散模型中的泛化与记忆现象。实证研究表明,这些现象受模型复杂度和训练数据规模影响。在实验中,我们进一步观察到,在去噪得分匹配(DSM)中,每个数据样本使用的噪声样本数 $m$ 起到显著且非平凡的作用。通过在一个简化理论框架下推导出 DSM 的测试误差和训练误差的渐近精确表达式,我们捕捉到了这些行为,并揭示了其机制。得分函数由随机特征神经网络参数化,目标分布为 $d$-维高斯分布。我们在一个极限情形下分析:维度 $d$、数据样本数 $n$、特征数 $p$ 同时趋于无穷,但比值 $ψ_n = n/d$ 与 $ψ_p = p/d$ 保持恒定。通过刻画测试误差与训练误差,我们识别出泛化与记忆的区域,其依赖于 $ψ_n, ψ_p$ 和 $m$。理论结果与实证观察一致。

原文摘要 · Abstract (English)

We theoretically investigate the phenomena of generalization and memorization in diffusion models. Empirical studies suggest that these phenomena are influenced by model complexity and the size of the training dataset. In our experiments, we further observe that the number of noise samples per data sample ($m$) used during Denoising Score Matching (DSM) plays a significant and non-trivial role. We capture these behaviors and shed insights into their mechanisms by deriving asymptotically precise expressions for test and train errors of DSM under a simple theoretical setting. The score function is parameterized by random features neural networks, with the target distribution being $d$-dimensional Gaussian. We operate in a regime where the dimension $d$, number of data samples $n$, and number of features $p$ tend to infinity while keeping the ratios $ψ_n=\frac{n}{d}$ and $ψ_p=\frac{p}{d}$ fixed. By characterizing the test and train errors, we identify regimes of generalization and memorization as a function of $ψ_n,ψ_p$, and $m$. Our theoretical findings are consistent with the empirical observations.

扩散模型得分匹配泛化分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。