arXiv:2511.08281cs.LG2025-11AAAI

重新审视重训练评估框架下的解释质量评价问题

Rethinking Explanation Evaluation under the Retraining Scheme

  • 通过重训练缓解输入扰动导致的分布偏移问题
  • 发现符号错误是造成评估偏差的关键因素
  • 提出新评估方法提升效率与可靠性,适合模型可解释性研究者

特征归因已成为解释模型决策的重要工具,但解释质量评估仍面临缺乏真实解释标签的挑战。为克服这一问题,基于推理的解释引导输入修改策略被广泛采用,通过观察输入变化对模型输出的影响来间接评估解释有效性。然而,此类方法常因扰动引入的数据分布偏移而影响评估可靠性。重训练方案ROAR通过调整模型以适应新分布来解决此问题,但其评估结果常与主流解释方法的理论预期相悖。本文深入分析该矛盾,识别出符号问题(sign issue)是导致残余信息干扰评估的核心原因。基于此,我们提出重构评估流程的简单方法,有效解决该问题。在此基础上,进一步设计新型变体,在保持原有框架基础上显著提升评估效率,增强解释器选择与基准测试的实际可行性。在多尺度数据集上的实验证明,新方法揭示了精心挑选解释器的性能特征,为可解释性研究中的开放挑战与未来方向提供了新洞见。

原文摘要 · Abstract (English)

Feature attribution has gained prominence as a tool for explaining model decisions, yet evaluating explanation quality remains challenging due to the absence of ground-truth explanations. To circumvent this, explanation-guided input manipulation has emerged as an indirect evaluation strategy, measuring explanation effectiveness through the impact of input modifications on model outcomes during inference. Despite the widespread use, a major concern with inference-based schemes is the distribution shift caused by such manipulations, which undermines the reliability of their assessments. The retraining-based scheme ROAR overcomes this issue by adapting the model to the altered data distribution. However, its evaluation results often contradict the theoretical foundations of widely accepted explainers. This work investigates this misalignment between empirical observations and theoretical expectations. In particular, we identify the sign issue as a key factor responsible for residual information that ultimately distorts retraining-based evaluation. Based on the analysis, we show that a straightforward reframing of the evaluation process can effectively resolve the identified issue. Building on the existing framework, we further propose novel variants that jointly structure a comprehensive perspective on explanation evaluation. These variants largely improve evaluation efficiency over the standard retraining protocol, thereby enhancing practical applicability for explainer selection and benchmarking. Following our proposed schemes, empirical results across various data scales provide deeper insights into the performance of carefully selected explainers, revealing open challenges and future directions in explainability research.

可解释性评估方法模型解释重训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。