arXiv:2502.07480cs.LGmath.ST2025-02NeurIPS被引 2

研究经典插值方法在噪声数据下的泛化行为,发现其过拟合可从灾难性到良性转变。

Beyond Benign Overfitting in Nadaraya-Watson Interpolators

  • 通过调节带宽超参数,揭示插值器存在多种过拟合模式。
  • 过拟合行为随参数变化非单调,经历灾难性、良性到温和的转变。
  • 高估数据内在维度比低估更安全,适合实际调参参考。

近年来,学界对插值预测器在含噪训练数据上的泛化行为产生浓厚兴趣。尽管标准分析关注方法是否一致,但近期观察表明,即使不一致的预测器也能良好泛化。本文重新审视经典的插值型Nadaraya-Watson(NW)估计器(又称Shepard法),从现代视角研究其泛化能力。通过调节单一类似带宽的超参数,我们证明了过拟合行为存在多重模式,且非单调地从灾难性、经由良性,过渡到温和状态。结果表明,即使是经典插值方法也表现出复杂的泛化特性。此外,为超参数调优提供启示:高估数据内在维度比低估更不具危害性。数值实验验证了上述理论现象。

原文摘要 · Abstract (English)

In recent years, there has been much interest in understanding the generalization behavior of interpolating predictors, which overfit on noisy training data. Whereas standard analyses are concerned with whether a method is consistent or not, recent observations have shown that even inconsistent predictors can generalize well. In this work, we revisit the classic interpolating Nadaraya-Watson (NW) estimator (also known as Shepard's method), and study its generalization capabilities through this modern viewpoint. In particular, by varying a single bandwidth-like hyperparameter, we prove the existence of multiple overfitting behaviors, ranging non-monotonically from catastrophic, through benign, to tempered. Our results highlight how even classical interpolating methods can exhibit intricate generalization behaviors. In addition, for the purpose of tuning the hyperparameter, the results suggest that over-estimating the intrinsic dimension of the data is less harmful than under-estimating it. Numerical experiments complement our theory, demonstrating the same phenomena.

插值器泛化分析过拟合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。