用逻辑数据集测试显著性方法,发现其能编码分类信息。
Saliency Methods are Encoders: Analysing Logical Relations Towards Interpretation
- 构建逻辑数据集,控制模型推理路径验证显著性方法
- 不同方法在互补与冗余信息下表现不一,但均能编码分类信息
- 适合关注可解释性评估的开发者与研究者
随着神经网络性能提升,架构日益复杂,解释性需求增强。当前大量新方法生成显著性图以提升可解释性,但评估常依赖视觉直觉,易产生确认偏误。由于缺乏通用解释质量度量、模型推理真实标签不可得,以及假设过多,诸多研究声称发现方法缺陷,却常导致不公平比较。此外,图像或文本等高复杂度数据集使所有可能解释难以逼近。为此,本文提出基于多重简单逻辑数据集的可控实验测试方法,通过分析模型在不同判别场景(如互补、冗余信息)下的推理关系,评估显著性方法对信息处理方式。引入多个新指标,以非信息性归因得分作为基线,检验典型期望偏差。结果表明,显著性方法能将分类相关的信息编码于显著性分数的排序之中。
原文摘要 · Abstract (English)
With their increase in performance, neural network architectures also become more complex, necessitating explainability. Therefore, many new and improved methods are currently emerging, which often generate so-called saliency maps in order to improve interpretability. Those methods are often evaluated by visual expectations, yet this typically leads towards a confirmation bias. Due to a lack of a general metric for explanation quality, non-accessible ground truth data about the model's reasoning and the large amount of involved assumptions, multiple works claim to find flaws in those methods. However, this often leads to unfair comparison metrics. Additionally, the complexity of most datasets (mostly images or text) is often so high, that approximating all possible explanations is not feasible. For those reasons, this paper introduces a test for saliency map evaluation: proposing controlled experiments based on all possible model reasonings over multiple simple logical datasets. Using the contained logical relationships, we aim to understand how different saliency methods treat information in different class discriminative scenarios (e.g. via complementary and redundant information). By introducing multiple new metrics, we analyse propositional logical patterns towards a non-informative attribution score baseline to find deviations of typical expectations. Our results show that saliency methods can encode classification relevant information into the ordering of saliency scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。