在重尾输入下,最大间隔分类器可实现无害过拟合,逼近理论最优误差。
Benign Overfitting under Learning Rate Conditions for $α$ Sub-exponential Input
- 研究重尾分布下的线性分类器,基于梯度下降训练
- 当维度与类别中心距满足条件时,错误率趋近噪声水平
- 发现学习率上限随尾部变重而降低,适用于真实数据场景
本文研究了具有重尾输入分布的二分类问题中的无害过拟合现象,将最大间隔分类器分析扩展至α次指数分布(α∈(0,2]),推广了以往仅针对次高斯输入的研究。在无正则化逻辑损失下,对使用梯度下降训练的线性分类器给出泛化误差界。结果表明,在特定维度p与类别中心间距条件下,最大间隔分类器的误分类误差渐近趋近于噪声水平,即理论最优值。同时推导出无害过拟合发生的学习率β上界,并证明随着输入分布尾部加重(α增大),该上界随之减小。这些结果表明,即使在比此前研究更重尾的输入设置下,无害过拟合仍可存在,深化了对更真实数据环境中该现象的理解。
原文摘要 · Abstract (English)
This paper investigates the phenomenon of benign overfitting in binary classification problems with heavy-tailed input distributions, extending the analysis of maximum margin classifiers to $α$ sub-exponential distributions ($α\in (0, 2]$). This generalizes previous work focused on sub-gaussian inputs. We provide generalization error bounds for linear classifiers trained using gradient descent on unregularized logistic loss in this heavy-tailed setting. Our results show that, under certain conditions on the dimensionality $p$ and the distance between the centers of the distributions, the misclassification error of the maximum margin classifier asymptotically approaches the noise level, the theoretical optimal value. Moreover, we derive an upper bound on the learning rate $β$ for benign overfitting to occur and show that as the tail heaviness of the input distribution $α$ increases, the upper bound on the learning rate decreases. These results demonstrate that benign overfitting persists even in settings with heavier-tailed inputs than previously studied, contributing to a deeper understanding of the phenomenon in more realistic data environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。