arXiv:2512.06347cs.LGstat.ML2025-12

随机插值器在样本足够多时可实现零泛化误差

Zero Generalization Error Theorem for Random Interpolators via Algebraic Geometry

  • 用代数几何分析参数空间中插值解的几何结构
  • 当样本数超过阈值时,随机插值器泛化误差为0
  • 适用于理解大模型泛化能力的理论研究

我们从理论上证明,在教师-学生框架下,当训练样本数量超过由参数空间中插值解集合几何结构决定的阈值后,机器学习模型的插值器将实现零泛化误差。尽管近年来理论研究认为深度神经网络的高泛化能力源于随机梯度下降(SGD)的隐式偏差,但实证证据表明,这主要源于模型自身的性质。具体而言,即使随机采样的插值器(即达到零训练误差的参数)也表现出良好的泛化性能。本研究通过代数几何工具,数学刻画了参数空间中插值解集合的几何结构,并证明当样本数超过该阈值时,随机插值器的泛化误差恰好为零。

原文摘要 · Abstract (English)

We theoretically demonstrate that the generalization error of interpolators for machine learning models under teacher-student settings becomes 0 once the number of training samples exceeds a certain threshold. Understanding the high generalization ability of large-scale models such as deep neural networks (DNNs) remains one of the central open problems in machine learning theory. While recent theoretical studies have attributed this phenomenon to the implicit bias of stochastic gradient descent (SGD) toward well-generalizing solutions, empirical evidences indicate that it primarily stems from properties of the model itself. Specifically, even randomly sampled interpolators, which are parameters that achieve zero training error, have been observed to generalize effectively. In this study, under a teacher-student framework, we prove that the generalization error of randomly sampled interpolators becomes exactly zero once the number of training samples exceeds a threshold determined by the geometric structure of the interpolator set in parameter space. As a proof technique, we leverage tools from algebraic geometry to mathematically characterize this geometric structure.

泛化误差插值器代数几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。