arXiv:2602.00825stat.MLcs.LG2026-02

在高维空间中,完美拟合噪声数据的函数会因平滑性偏差导致泛化误差无法消失。

Harmful Overfitting in Sobolev Spaces

  • 利用索博列夫空间中的范数最小化插值器,研究过拟合机制
  • 即使样本量无穷大,泛化误差仍被正数下界限制
  • 突破传统希尔伯特空间限制,适用于任意p值

受过参数化机器学习中良性过拟合研究的启发,本文研究了在索博列夫空间 $W^{k, p}(bR^d)$ 中完美拟合含噪声训练数据的函数的泛化行为。在标签噪声和数据分布足够光滑的假设下,我们证明:近似范数最小化的插值器(由平滑性偏好选择的典型解)表现出有害过拟合——即使训练样本数 $n \to \infty$,泛化误差仍以高概率被一个正数下界所限制。该结果对任意 $p \in [1, \infty)$ 成立,不同于以往仅针对希尔伯特空间($p = 2$)且基于核方法的研究。证明采用几何论证,通过索博列夫不等式识别出训练数据附近的有害邻域。

原文摘要 · Abstract (English)

Motivated by recent work on benign overfitting in overparameterized machine learning, we study the generalization behavior of functions in Sobolev spaces $W^{k, p}(\mathbb{R}^d)$ that perfectly fit a noisy training data set. Under assumptions of label noise and sufficient regularity in the data distribution, we show that approximately norm-minimizing interpolators, which are canonical solutions selected by smoothness bias, exhibit harmful overfitting: even as the training sample size $n \to \infty$, the generalization error remains bounded below by a positive constant with high probability. Our results hold for arbitrary values of $p \in [1, \infty)$, in contrast to prior results studying the Hilbert space case ($p = 2$) using kernel methods. Our proof uses a geometric argument which identifies harmful neighborhoods of the training data using Sobolev inequalities.

泛化理论过拟合索博列夫空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。