arXiv:2501.14136cs.LG2025-01被引 1

揭示了显著性图在逻辑推理任务中的根本性缺陷。

Saliency Maps are Ambiguous: Analysis of Logical Relations on First and Second Order Attributions

  • 通过扩展ANDOR框架,分析多种数据集上的显著性方法
  • 发现现有方法无法准确捕捉所有分类所需信息
  • 提出全局一致性表示法,实现真正输入忽略的评估

近期研究揭示了基于归因或热力图的显著性方法可能存在缺陷。典型问题包括确认偏差——评分与人类预期对比。由于缺乏模型推理的真值,评估显著性方法质量困难,发现普遍局限也难以实现。这进一步复杂化,因为对复杂数据的掩码评估易引入偏差,多数方法无法完全忽略输入。本文扩展先前在逻辑数据集框架ANDOR上的分析,表明所有测试的显著性方法在所有可能场景中均未能掌握必要的分类信息。本文通过增加更多数据集分析,更好理解方法失效的场景;并引入全局一致性表示作为额外评估方法,以实现真正的输入删除评估。

原文摘要 · Abstract (English)

Recent work uncovered potential flaws in \eg attribution or heatmap based saliency methods. A typical flaw is a confirmations bias, where the scores are compared to human expectation. Since measuring the quality of saliency methods is hard due to missing ground truth model reasoning, finding general limitations is also hard. This is further complicated, because masking-based evaluation on complex data can easily introduce a bias, as most methods cannot fully ignore inputs. In this work, we extend our previous analysis on the logical dataset framework ANDOR, where we showed that all analysed saliency methods fail to grasp all needed classification information for all possible scenarios. Specifically, this paper extends our previous work using analysis on more datasets, in order to better understand in which scenarios the saliency methods fail. Further, we apply the Global Coherence Representation as an additional evaluation method in order to enable actual input omission.

显著性分析归因方法逻辑推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。