揭示生成模型记忆与过拟合的理论机制,解释样本为何接近训练数据。
A Theoretical Analysis of Memory and Overfitting Phenomena in Stochastic Interpolation Models

- 基于闭式解分析最优速度场与得分函数,推导生成过程本质
- 离散化误差与估计误差共同控制生成样本偏离训练集程度
- 提出生成模型过拟合/欠拟合的理论定义,适合理论研究者
本文提供了对随机插值生成模型中记忆现象的理论分析。通过利用最优速度场和相关得分函数的闭式表达式,我们证明在连续时间理想设定下,确定性和随机生成过程均可恢复训练样本。在欧拉离散化下,生成样本仍围绕训练样本分布,偏差由步长控制。进一步分析估计误差下的生成过程表明,累积估计误差决定了终点偏离训练集的程度。这些结果表明,生成样本可表征为训练样本受三类可控项扰动的结果:离散化引起的有界偏差、估计误差引起的有界偏差及随机高斯噪声。基于此表征,本文给出了生成模型过拟合与欠拟合的理论定义。合成仿真验证了理论结论。
原文摘要 · Abstract (English)
This paper provides a theoretical account of memorization in stochastic interpolation models. By leveraging closed-form expressions for the optimal velocity field and the associated score function, we show that, in the continuous-time oracle setting, both deterministic and stochastic generation processes recover training samples. Under Euler discretization, generated samples remain centered around training samples, with deviations controlled by the step size. We further analyze generation in the presence of estimation errors and show that accumulated estimation errors control the endpoint deviation from the training set. These results imply that the generated sample admits a representation as a training sample perturbed by three controlled terms: a discretization-induced bound, an estimation-error-induced bound, and stochastic Gaussian noise. Based on this characterization, we provide theoretical definitions of overfitting and underfitting in generative models. Synthetic simulations support our theoretical findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。