arXiv:2501.00961cs.LGcs.AI2025-01中稿 · Nature Communicati…被引 26

发现少数群体信息被少量神经元过度记忆,导致模型测试时表现失衡。

Uncovering Memorization Effect in the Presence of Spurious Correlations

  • 通过分析神经元激活,发现少数群体特征被集中在少数神经元中记忆。
  • 在多个模型和数据集上,移除这些虚假记忆可显著提升少数群体准确率。
  • 研究揭示了模型对无关特征的过度记忆是性能不均的关键原因,适合关注公平性与鲁棒性的研究者。

机器学习模型常依赖训练数据中的简单虚假特征(如图像背景与前景分类的相关性),这类特征虽与目标相关但无因果关系,导致少数群体在测试时表现显著低于多数群体。本文从记忆化角度深入分析这一现象,首次系统证明虚假特征存在于网络中极少数神经元内。通过三种实验路径的交叉验证,发现少量神经元或通道会记忆少数群体信息。基于此,我们提出假设:集中于少数神经元的虚假记忆是造成群体性能不平衡的关键因素。进一步提出一种新型训练框架,在训练阶段消除这些不必要的虚假记忆模式,实验证明该方法能显著改善模型在少数群体上的表现。结果在多种架构与基准上一致有效,揭示了神经网络对核心知识与虚假知识的编码机制,为未来研究模型对虚假相关性的鲁棒性提供新思路。

原文摘要 · Abstract (English)

Machine learning models often rely on simple spurious features -- patterns in training data that correlate with targets but are not causally related to them, like image backgrounds in foreground classification. This reliance typically leads to imbalanced test performance across minority and majority groups. In this work, we take a closer look at the fundamental cause of such imbalanced performance through the lens of memorization, which refers to the ability to predict accurately on atypical examples (minority groups) in the training set but failing in achieving the same accuracy in the testing set. This paper systematically shows the ubiquitous existence of spurious features in a small set of neurons within the network, providing the first-ever evidence that memorization may contribute to imbalanced group performance. Through three experimental sources of converging empirical evidence, we find the property of a small subset of neurons or channels in memorizing minority group information. Inspired by these findings, we hypothesize that spurious memorization, concentrated within a small subset of neurons, plays a key role in driving imbalanced group performance. To further substantiate this hypothesis, we show that eliminating these unnecessary spurious memorization patterns via a novel framework during training can significantly affect the model performance on minority groups. Our experimental results across various architectures and benchmarks offer new insights on how neural networks encode core and spurious knowledge, laying the groundwork for future research in demystifying robustness to spurious correlation.

模型公平性虚假相关神经元记忆性能失衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。