arXiv:2505.10399cs.AIcs.LG2025-05被引 10

无需真实解释即可评估模型解释质量,解决解释可信度判定难题。

Evaluating Model Explanations without Ground Truth

  • 提出无真值依赖的解释评估框架AXE,不需理想解释或模型敏感性验证。
  • 在多个数据集上验证其有效性,能准确区分真实与虚假解释。
  • 适合研究解释公平性、对抗攻击检测及模型可解释性评估的学者使用。

单一模型预测可能存在多种相互矛盾的解释,导致难以选择可信解释。现有评估方法依赖理想‘真值’解释或模型对重要输入的敏感性验证,但存在局限性。本文指出这些方法的问题,提出三个评估解释质量应遵循的原则,并构建无需真值的解释评估框架AXE。该框架不依赖理想解释或模型敏感性,提供独立的解释质量衡量标准。通过与基线对比验证了其有效性,展示了其在识别解释‘漂白’(fairwashing)中的应用潜力。代码已开源。

原文摘要 · Abstract (English)

There can be many competing and contradictory explanations for a single model prediction, making it difficult to select which one to use. Current explanation evaluation frameworks measure quality by comparing against ideal "ground-truth" explanations, or by verifying model sensitivity to important inputs. We outline the limitations of these approaches, and propose three desirable principles to ground the future development of explanation evaluation strategies for local feature importance explanations. We propose a ground-truth Agnostic eXplanation Evaluation framework (AXE) for evaluating and comparing model explanations that satisfies these principles. Unlike prior approaches, AXE does not require access to ideal ground-truth explanations for comparison, or rely on model sensitivity - providing an independent measure of explanation quality. We verify AXE by comparing with baselines, and show how it can be used to detect explanation fairwashing. Our code is available at https://github.com/KaiRawal/Evaluating-Model-Explanations-without-Ground-Truth.

可解释性模型评估公平性检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。