arXiv:2509.04588cs.LGcs.AI2025-09

提出新方法让解释同时保持计算路径一致性与输出准确性。

Beyond Output Faithfulness: Learning Attributions that Preserve Computational Pathways

  • 用激活保留作为计算路径的可操作代理,避免解释偏离真实计算过程。
  • 在ImageNet和CUB-200-2011上实现最优插入/删除得分,激活偏差降低40%以上。
  • 适合追求高可信度解释的模型可解释性研究者使用。

现有忠实性度量如插入/删除法评估特征移除对输出的影响,但忽略了解释是否保留网络实际使用的计算路径。我们发现,外部度量可通过替代路径最大化——即通过不同特征检测器重定向计算,同时维持输出行为不变。为此,我们提出激活保留作为计算路径保留的可行代理,并引入基于忠实性的集成解释(FEI)方法,联合优化外部忠实性(通过插入/删除曲线的集成分位数优化)与内部忠实性(通过选择性梯度裁剪)。在VGG和ResNet于ImageNet与CUB-200-2011上的实验表明,FEI在取得当前最优插入/删除得分的同时,激活偏差显著更低,证明外部与内部忠实性对可靠解释均至关重要。

原文摘要 · Abstract (English)

Faithfulness metrics such as insertion and deletion evaluate how feature removal affects model outputs but overlook whether explanations preserve the computational pathway the network actually uses. We show that external metrics can be maximized through alternative pathways -- perturbations that reroute computation via different feature detectors while preserving output behavior. To address this, we propose activation preservation as a tractable proxy for preserving computational pathways We introduce Faithfulness-guided Ensemble Interpretation (FEI), which jointly optimizes external faithfulness (via ensemble quantile optimization of insertion/deletion curves) and internal faithfulness (via selective gradient clipping). Across VGG and ResNet on ImageNet and CUB-200-2011, FEI achieves state-of-the-art insertion/deletion scores while maintaining significantly lower activation deviation, showing that both external and internal faithfulness are essential for reliable explanations.

可解释性模型忠实性激活保留

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。