arXiv:2510.06401cs.LGcs.IT2025-10被引 3

研究标签噪声对神经网络表示信息量的影响,发现过参数模型对噪声鲁棒。

The Effect of Label Noise on the Information Content of Neural Representations

  • 用信息失衡度衡量隐藏层表示的信息含量,分析噪声标签下的表示变化。
  • 在过参数化阶段,噪声标签训练的表示与干净标签相当,信息量不降反稳。
  • 揭示过参数模型在噪声下仍能保持良好泛化,适合高噪声数据场景应用。

在监督分类任务中,模型需为每个数据点预测标签。真实数据集中的标签常因标注错误而存在噪声。尽管标签噪声对深度学习模型性能的影响已广受关注,但其对网络隐含表示的影响仍不清楚。本文通过信息失衡(Information Imbalance)这一计算高效的条件互信息代理,系统比较了不同条件下隐藏表示的信息含量。结果发现,隐藏表示的信息量随网络参数数量呈双下降趋势,与测试误差行为类似。在参数不足阶段,使用噪声标签训练的表示比使用干净标签更具有信息性;而在参数充足阶段,两者信息量相当。这表明过参数化网络的表示对标签噪声具有鲁棒性。此外,在过参数化阶段,倒数第二层与预softmax层间的信息失衡随交叉熵损失减小而降低,为理解分类任务的泛化提供了新视角。进一步分析随机标签学习的表示发现,其表现劣于随机特征,说明随机标签训练使网络远超懒惰学习状态,权重开始主动编码标签信息。

原文摘要 · Abstract (English)

In supervised classification tasks, models are trained to predict a label for each data point. In real-world datasets, these labels are often noisy due to annotation errors. While the impact of label noise on the performance of deep learning models has been widely studied, its effects on the networks' hidden representations remain poorly understood. We address this gap by systematically comparing hidden representations using the Information Imbalance, a computationally efficient proxy of conditional mutual information. Through this analysis, we observe that the information content of the hidden representations follows a double descent as a function of the number of network parameters, akin to the behavior of the test error. We further demonstrate that in the underparameterized regime, representations learned with noisy labels are more informative than those learned with clean labels, while in the overparameterized regime, these representations are equally informative. Our results indicate that the representations of overparameterized networks are robust to label noise. We also found that the information imbalance between the penultimate and pre-softmax layers decreases with cross-entropy loss in the overparameterized regime. This offers a new perspective on understanding generalization in classification tasks. Extending our analysis to representations learned from random labels, we show that these perform worse than random features. This indicates that training on random labels drives networks much beyond lazy learning, as weights adapt to encode labels information.

表示学习标签噪声过参数化信息理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。