揭示扩散模型在高维下泛化、记忆与过拟合的三阶段机制
Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime
- 通过核方法分析梯度流训练下的风险轨迹,发现三类不同估计器
- 在高维比例极限下,模型经历泛化、插值和记忆三个阶段
- 适用于研究生成模型理论性质的研究者,特别是扩散模型方向
现代基于得分的生成模型在图像、音频和视频合成等高维任务中取得了显著成功。这些模型将分布学习转化为一系列回归问题,若在有限数据上精确求解,最终会复现训练样本。其泛化能力必然源于训练过程中的隐式或显式正则化。本文发展了针对过参数神经网络在懒惰训练区间的良性过拟合与算法正则化理论的生成模型对应版本。研究在向量值再生核希尔伯特空间中使用内积核的去噪得分匹配。在比例高维极限 $n/asymptotic d$ 下,推导出梯度流训练下的精确风险轨迹。这些轨迹表现出由三类定性不同的估计器主导的三个阶段:一个能泛化的谱估计器、一个具有局部峰值的纯噪声得分(插值训练目标)、以及一个记忆数据的经验贝叶斯估计器。随后分析这些估计器在反向时间SDE中的组合方式,并刻画生成样本的分布。分析揭示了监督学习中的熟悉机制,如核线性化及核非线性部分带来的自诱导正则化,但也暴露出生成建模特有的独特现象。
原文摘要 · Abstract (English)
Modern score-based generative models have achieved remarkable empirical success in high-dimensional tasks such as image, audio, and video synthesis. These models reduce distribution learning to a sequence of regression problems that, if solved exactly on finite data, would ultimately reproduce the training samples. Their ability to generalize must therefore arise from the implicit or explicit regularization during training. In this work, we develop a generative counterpart to the theory of benign overfitting and algorithmic regularization for overparameterized neural networks in the supervised lazy-training regime. We study denoising score matching in a vector-valued reproducing kernel Hilbert space with an inner-product kernel. In the proportional high-dimensional regime $n\asymp d$, we derive exact risk trajectories under gradient flow training. These trajectories exhibit three phases governed by qualitatively distinct estimators: a spectral estimator that generalizes, a pure-noise score with localized peaks that interpolate the training objective, and an empirical Bayes estimator that memorizes the data. We then analyze how these estimators combine along the reverse-time SDE and characterize the distribution of the resulting samples. The analysis reveals familiar mechanisms from supervised learning, including kernel linearization and self-induced regularization from the nonlinear part of the kernel, but also reveals a distinct phenomenology specific to generative modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。