新模型揭示神经网络如何利用噪声提升长尾数据分类效果
Rethinking Benign Overfitting in Two-Layer Neural Networks
- 引入类别相关噪声,改进传统过拟合理论模型
- 证明神经网络可借噪声学习隐式特征,降低测试误差
- 提出无需训练的数据影响评估指标,适合长尾分布研究
近期理论研究(Kou et al., 2023;Cao et al., 2022)发现,当噪声与特征比超过阈值时,过拟合会从良性转为有害——这在长尾数据分布中常见。然而,实际中过参数化神经网络很少出现有害过拟合。实验也表明,在长尾分布下,记忆数据对实现近最优泛化误差是必要的(Feldman & Zhang, 2020)。我们指出,这一理论与实证的差异源于以往模型忽略了不同类别间噪声的异质性。本文通过引入类别依赖的异质噪声,重构特征-噪声数据模型,并系统分析训练动态,推导出修正后模型的测试损失上界。结果表明,神经网络能利用‘数据噪声’学习隐式特征,从而提升长尾数据分类准确率。此外,我们的分析还提供了无需训练即可评估数据对测试性能影响的度量方法。在合成与真实数据集上的实验验证了理论结论。
原文摘要 · Abstract (English)
Recent theoretical studies (Kou et al., 2023; Cao et al., 2022) have revealed a sharp phase transition from benign to harmful overfitting when the noise-to-feature ratio exceeds a threshold-a situation common in long-tailed data distributions where atypical data is prevalent. However, harmful overfitting rarely happens in overparameterized neural networks. Further experimental results suggested that memorization is necessary for achieving near-optimal generalization error in long-tailed data distributions (Feldman & Zhang, 2020). We argue that this discrepancy between theoretical predictions and empirical observations arises because previous feature-noise data models overlook the heterogeneous nature of noise across different data classes. In this paper, we refine the feature-noise data model by incorporating class-dependent heterogeneous noise and re-examine the overfitting phenomenon in neural networks. Through a comprehensive analysis of the training dynamics, we establish test loss bounds for the refined model. Our findings reveal that neural networks can leverage "data noise" to learn implicit features that improve the classification accuracy for long-tailed data. Our analysis also provides a training-free metric for evaluating data influence on test performance. Experimental validation on both synthetic and real-world datasets supports our theoretical results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。