提出新框架,让可视化解释更贴合人类提问方式。
Rethinking Saliency Maps: A Cognitive Human Aligned Taxonomy and Evaluation Framework for Explanations
- 按'谁问的'和'回答多细'将解释分类,区分点对点与对比性问题。
- 发现现有评估方法只重准确性,忽略对比推理和语义层次。
- 设计四种新指标,可全面评测解释是否真正回应用户疑问。
Saliency maps 广泛用于深度学习的视觉解释,但其目的和与用户查询的对齐仍缺乏共识,影响了评估与实际应用。本文提出参考帧×粒度(RFxG)分类体系,从两个维度组织解释:参考帧(点对点:“为何预测为此?”与对比性:“为何是这个而非其他?”)和粒度(细粒度类别级如“为何是哈士奇?”至粗粒度组级如“为何是狗?”)。基于此,我们揭示现有评估指标普遍偏重点对点忠实性,忽视对比推理与语义粒度。为此,提出四项新忠实性指标,系统评估十种前沿显著性方法、四种模型架构及三个数据集。本研究倡导以用户意图为导向的评估,为开发既忠实于模型行为又契合人类认知复杂性的视觉解释提供了理论基础与实用工具。
原文摘要 · Abstract (English)
Saliency maps are widely used for visual explanations in deep learning, but a fundamental lack of consensus persists regarding their intended purpose and alignment with diverse user queries. This ambiguity hinders the effective evaluation and practical utility of explanation methods. We address this gap by introducing the Reference-Frame $\times$ Granularity (RFxG) taxonomy, a principled conceptual framework that organizes saliency explanations along two essential axes:Reference-Frame: Distinguishing between pointwise ("Why this prediction?") and contrastive ("Why this and not an alternative?") explanations. Granularity: Ranging from fine-grained class-level (e.g., "Why Husky?") to coarse-grained group-level (e.g., "Why Dog?") interpretations. Using the RFxG lens, we demonstrate critical limitations in existing evaluation metrics, which overwhelmingly prioritize pointwise faithfulness while neglecting contrastive reasoning and semantic granularity. To systematically assess explanation quality across both RFxG dimensions, we propose four novel faithfulness metrics. Our comprehensive evaluation framework applies these metrics to ten state-of-the-art saliency methods, four model architectures, and three datasets. By advocating a shift toward user-intent-driven evaluation, our work provides both the conceptual foundation and the practical tools necessary to develop visual explanations that are not only faithful to the underlying model behavior but are also meaningfully aligned with the complexity of human understanding and inquiry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。