有限训练样本下,生成模型的随机插值行为可被精确刻画。
Generation Properties of Stochastic Interpolation under Finite Training Set
- 推导出有限样本下的最优速度场与得分函数闭式解。
- 随机生成过程等价于训练样本加高斯噪声,确定性过程可完全复现训练数据。
- 揭示生成模型过拟合与欠拟合的新定义,适用于实际建模场景。
本文研究了在有限训练样本条件下生成模型的理论行为。在随机插值生成框架下,我们推导出仅使用有限训练样本时最优速度场与得分函数的闭式表达式。在某些正则条件下,确定性生成过程能精确恢复训练样本,而随机生成过程表现为带高斯噪声的训练样本。在非理想设置中,考虑模型估计误差,提出生成模型特有的过拟合与欠拟合形式化定义。理论分析表明,在存在估计误差时,随机生成过程等效于训练样本的凸组合,且受均匀噪声与高斯噪声混合污染。生成任务及下游分类任务的实验验证了理论结果。
原文摘要 · Abstract (English)
This paper investigates the theoretical behavior of generative models under finite training populations. Within the stochastic interpolation generative framework, we derive closed-form expressions for the optimal velocity field and score function when only a finite number of training samples are available. We demonstrate that, under some regularity conditions, the deterministic generative process exactly recovers the training samples, while the stochastic generative process manifests as training samples with added Gaussian noise. Beyond the idealized setting, we consider model estimation errors and introduce formal definitions of underfitting and overfitting specific to generative models. Our theoretical analysis reveals that, in the presence of estimation errors, the stochastic generation process effectively produces convex combinations of training samples corrupted by a mixture of uniform and Gaussian noise. Experiments on generation tasks and downstream tasks such as classification support our theory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。