arXiv:2504.15664cs.LGcs.CV2025-04中稿 · The World Conferen…被引 2

用可解释AI分析神经网络如何依赖虚假特征,发现不同模型差异。

An XAI-based Analysis of Shortcut Learning in Neural Networks

  • 提出神经元虚假得分,量化神经元对虚假特征的依赖程度。
  • 发现卷积网络和视觉变压器对虚假特征的解耦程度不同。
  • 揭示现有缓解方法假设不完整,为新方法提供基础。

机器学习模型容易学习虚假特征——与目标标签强相关但非因果的特征。现有缓解模型对虚假特征依赖的方法在某些情况下有效,但在其他情况下失败。本文系统分析了神经网络如何以及在何处编码虚假相关性。我们引入基于可解释AI的神经元虚假得分,用于量化神经元对虚假特征的依赖程度。通过针对架构特异的方法,分析了卷积神经网络(CNNs)和视觉变换器(ViTs)。结果表明,虚假特征部分被解耦,但解耦程度随模型架构而异。此外,我们发现现有缓解方法背后的假设不完整。研究结果为开发新型缓解虚假相关性的方法奠定基础,使AI模型在实际应用中更安全。

原文摘要 · Abstract (English)

Machine learning models tend to learn spurious features - features that strongly correlate with target labels but are not causal. Existing approaches to mitigate models' dependence on spurious features work in some cases, but fail in others. In this paper, we systematically analyze how and where neural networks encode spurious correlations. We introduce the neuron spurious score, an XAI-based diagnostic measure to quantify a neuron's dependence on spurious features. We analyze both convolutional neural networks (CNNs) and vision transformers (ViTs) using architecture-specific methods. Our results show that spurious features are partially disentangled, but the degree of disentanglement varies across model architectures. Furthermore, we find that the assumptions behind existing mitigation methods are incomplete. Our results lay the groundwork for the development of novel methods to mitigate spurious correlations and make AI models safer to use in practice.

可解释AI虚假特征神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。