arXiv:2504.08553cs.LGcs.AI2025-04中稿 · XAI World Conferen…被引 1

用谱分析揭示解释质量的两个核心维度:稳定性和目标敏感性。

Uncovering the Structure of Explanation Quality with Spectral Analysis

  • 通过谱分解方法识别解释质量的关键特性。
  • 实验发现现有评估方法仅部分捕捉稳定性与敏感性的权衡。
  • 为改进解释评估提供可解释的理论框架,适合模型可解释性研究者。

随着机器学习模型在高风险领域应用日益广泛,有效的解释方法对于确保其预测策略对用户透明至关重要。多年来,已提出众多指标来评估解释质量,但其实际适用性仍不明确,尤其因对各指标所奖励的具体方面理解有限。本文提出一种基于解释结果谱分析的新框架,系统捕获不同解释技术的多维特性。分析揭示出解释质量的两个独立因素:稳定性与目标敏感性,可通过谱分解直接观察。在MNIST和ImageNet上的实验表明,主流评估方法(如像素翻转、熵)仅部分反映这两者的权衡。总体而言,该框架为理解解释质量提供了基础,有助于指导更可靠的解释评估技术发展。

原文摘要 · Abstract (English)

As machine learning models are increasingly considered for high-stakes domains, effective explanation methods are crucial to ensure that their prediction strategies are transparent to the user. Over the years, numerous metrics have been proposed to assess quality of explanations. However, their practical applicability remains unclear, in particular due to a limited understanding of which specific aspects each metric rewards. In this paper we propose a new framework based on spectral analysis of explanation outcomes to systematically capture the multifaceted properties of different explanation techniques. Our analysis uncovers two distinct factors of explanation quality-stability and target sensitivity-that can be directly observed through spectral decomposition. Experiments on both MNIST and ImageNet show that popular evaluation techniques (e.g., pixel-flipping, entropy) partially capture the trade-offs between these factors. Overall, our framework provides a foundational basis for understanding explanation quality, guiding the development of more reliable techniques for evaluating explanations.

可解释性谱分析模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。