arXiv:2608.26083cs.LGcs.AI2026-08

提出新方法精准识别神经网络中的虚假关联,避免误报。

ICON Decomposition: Auditing Deep Neural Networks with Multivariate Variance-based Concept-level Explanations

论文配图:ICON Decomposition: Auditing Deep Neural Networks with Multivariate Variance-based Concept-level Explanations
图 1 · 摘自论文原文
  • 基于多变量方差分解,量化每个概念对特征层的解释力
  • 在皮肤癌模型中成功检测出人为植入的捷径,基线方法误报
  • 适用于医学影像等高风险场景,帮助验证模型可靠性

深度神经网络常依赖虚假关联,即所谓的捷径学习。审计此类问题需测试多个候选概念,如采集条件或人口统计信息。现有基于概念的可解释性方法仅逐个测试概念是否可从某层解码,易将相关性误判为因果性,且不同层或概念类型间得分不可比。本文提出ICON分解,量化每个概念在已知其他概念与输出情况下的层内方差解释量,以及所有概念均无法解释的部分,生成层间可比、校准后的得分,有效抑制假阳性。在模拟数据上,ICON比七种现有方法更准确恢复概念重要性;在植入伪影的皮肤癌模型中,成功检测到诱导的捷径,而基线方法产生假阳性;在两个脑影像模型中,其稀疏解释经重训练和分布外测试验证有效。

原文摘要 · Abstract (English)

Deep neural networks often exploit spurious associations, a failure known as shortcut learning. Auditing for shortcuts requires testing many candidate concepts, such as acquisition settings or demographics. Current concept-based explainability methods test one concept at a time, asking whether it is decodable from a layer. These methods flag concepts that are merely correlated with the outcome or with each other, and their scores are not comparable across layers or concept types. We introduce ICON decomposition, which quantifies how much of a layer's variance each concept explains given all other concepts and the outcome, and how much none of them explains, yielding layer-comparable, calibrated scores that suppress false positives. On simulated data, ICON recovers concept importance more accurately than seven existing methods. On skin-cancer models with inserted artifacts, it detects induced shortcuts where baselines report false positives. On two brain-imaging models, ICON's sparse explanations are validated by retraining and out-of-distribution tests.

可解释性模型审计捷径学习方差分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。