arXiv:2503.07346cs.CVcs.LG2025-03被引 1

用多类别分布提升注意力图精度,让模型决策更透明。

Hidden in Plain Sight -- Class Competition Focuses Attribution Maps

  • 以多类概率分布代替单个输出作为归因目标
  • 在7种架构上使18种方法性能最高提升2倍
  • 无需改动模型,通用性强,适合想理解模型的开发者

归因方法揭示神经网络预测所依赖的输入特征,提升决策透明度。但常见问题是归因结果模糊,同时突出重要和无关特征。我们重新审视归因流程,发现使用logits作为归因目标是导致该现象的主要原因。解决方案就在眼前:利用现有归因方法对多个类别的归因分布进行分析,可获得更精确、细粒度的归因结果。在包括网格指向任务和基于随机化的合理性检验在内的多个基准测试中,该方法使7种架构上的18种归因方法性能最高提升2倍,且不依赖具体模型结构。

原文摘要 · Abstract (English)

Attribution methods reveal which input features a neural network uses for a prediction, adding transparency to their decisions. A common problem is that these attributions seem unspecific, highlighting both important and irrelevant features. We revisit the common attribution pipeline and observe that using logits as attribution target is a main cause of this phenomenon. We show that the solution is in plain sight: considering distributions of attributions over multiple classes using existing attribution methods yields specific and fine-grained attributions. On common benchmarks, including the grid-pointing game and randomization-based sanity checks, this improves the ability of 18 attribution methods across 7 architectures up to 2x, agnostic to model architecture.

模型可解释性归因方法注意力图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。