揭示线性分类器在噪声数据中仍能良好泛化的普遍规律
Universality of Benign Overfitting in Binary Linear Classification
- 通过几何分析发现噪声模型下测试误差的相变现象
- 放宽了对特征分布的强假设,扩展了适用场景范围
- 为深度学习过拟合却泛化良好的现象提供新解释
深度学习的实际成功催生了诸多意外现象。其中最受关注的是“良性过拟合”:尽管深度神经网络在含噪训练数据上达到完美拟合,仍能良好泛化。目前理论已证明该现象存在于多种经典统计模型中。对于线性最大间隔分类器,已有研究在特定混合模型下建立了理论结果,但前提是对协变量分布有极强假设。尤其在真实场景更常见的噪声标签情况下,理解仍不充分。本文对线性最大间隔分类器的良性过拟合进行了全面研究,首次揭示了噪声模型中测试误差界存在的相变现象,并给出相应几何解释。同时,在噪声与无噪声情形下均显著弱化了对协变量分布的要求。结果表明,最大间隔分类器的良性过拟合可存在于比以往认知更广泛的情境中,为内在机制提供了新见解。
原文摘要 · Abstract (English)
The practical success of deep learning has led to the discovery of several surprising phenomena. One of these phenomena, that has spurred intense theoretical research, is ``benign overfitting'': deep neural networks seem to generalize well in the over-parametrized regime even though the networks show a perfect fit to noisy training data. It is now known that benign overfitting also occurs in various classical statistical models. For linear maximum margin classifiers, benign overfitting has been established theoretically in a class of mixture models with very strong assumptions on the covariate distribution. However, even in this simple setting, many questions remain open. For instance, most of the existing literature focuses on the noiseless case where all true class labels are observed without errors, whereas the more interesting noisy case remains poorly understood. We provide a comprehensive study of benign overfitting for linear maximum margin classifiers. We discover a phase transition in test error bounds for the noisy model which was previously unknown and provide some geometric intuition behind it. We further considerably relax the required covariate assumptions in both the noisy and noiseless cases. Our results demonstrate that benign overfitting of maximum margin classifiers holds in a much wider range of scenarios than was previously known and provide new insights into the underlying mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。