arXiv:2511.10439cs.LGcs.AI2025-11NeurIPS被引 2

通过校准不确定性提升解释可靠性,让模型在扰动下更可信。

Improving Perturbation-based Explanations by Understanding the Role of Uncertainty Calibration

  • 用新方法ReCalX校准模型在扰动下的置信度
  • 实验显示能显著降低扰动带来的置信度偏差
  • 适合关注模型可解释性与稳定性的研究者

基于扰动的解释广泛用于提升机器学习模型的透明性,但其可靠性常因模型在特定扰动下的未知行为而受损。本文研究了不确定性校准(即模型置信度与实际准确率的一致性)与扰动解释之间的关系,发现模型在解释专用扰动下会产生系统性不可靠的概率估计,并从理论上证明这会直接损害全局与局部解释质量。为此,提出ReCalX,一种在不改变原始预测的前提下优化解释性能的新方法。跨多种模型与数据集的实证评估表明,ReCalX能最有效地减少扰动特定的校准偏差,同时提升解释鲁棒性,并更准确识别全局重要输入特征。

原文摘要 · Abstract (English)

Perturbation-based explanations are widely utilized to enhance the transparency of machine-learning models in practice. However, their reliability is often compromised by the unknown model behavior under the specific perturbations used. This paper investigates the relationship between uncertainty calibration - the alignment of model confidence with actual accuracy - and perturbation-based explanations. We show that models systematically produce unreliable probability estimates when subjected to explainability-specific perturbations and theoretically prove that this directly undermines global and local explanation quality. To address this, we introduce ReCalX, a novel approach to recalibrate models for improved explanations while preserving their original predictions. Empirical evaluations across diverse models and datasets demonstrate that ReCalX consistently reduces perturbation-specific miscalibration most effectively while enhancing explanation robustness and the identification of globally important input features.

可解释性不确定性校准扰动解释模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。