提出解释可靠性度量,检验XAI方法在真实场景下的稳定性。
Reliable Explanations or Random Noise? A Reliability Metric for XAI
- 设计四个可靠性公理,量化解释在扰动下的稳定性
- 实验证明主流解释方法在真实条件下常不稳定
- 提供基准测试工具,适合评估和改进XAI系统
近年来,解释复杂机器学习模型决策在能源系统、医疗、金融和自动驾驶等高风险领域变得至关重要。然而,这些解释的可靠性——即在现实非对抗性变化下是否保持稳定一致——仍缺乏有效度量。尽管SHAP和集成梯度(IG)等方法在归因理论上具有合理性,但其解释在小输入扰动、特征相关性及模型微调等系统级条件下可能显著变化。这种变异性损害了解释的可靠性,因为可靠解释应在等价输入表示和小幅性能保持的模型变化中保持一致。本文提出解释可靠性指数(ERI),一组基于四项可靠性公理的度量:对小输入扰动的鲁棒性、特征冗余下的一致性、模型演进中的平滑性以及温和分布偏移下的韧性。针对每项公理,推导出形式化保证,包括Lipschitz型边界与时间稳定性结果。我们进一步提出ERI-T,用于序列模型的时间可靠性度量,并引入ERI-Bench基准,系统性地在合成与真实数据集上压力测试解释可靠性。实验揭示主流解释方法普遍存在可靠性失效问题,表明解释在真实部署条件下可能极不稳定。通过暴露并量化这些不稳定性,ERI实现了对解释可靠性的原则性评估,推动更可信的可解释人工智能(XAI)系统发展。
原文摘要 · Abstract (English)
In recent years, explaining decisions made by complex machine learning models has become essential in high-stakes domains such as energy systems, healthcare, finance, and autonomous systems. However, the reliability of these explanations, namely, whether they remain stable and consistent under realistic, non-adversarial changes, remains largely unmeasured. Widely used methods such as SHAP and Integrated Gradients (IG) are well-motivated by axiomatic notions of attribution, yet their explanations can vary substantially even under system-level conditions, including small input perturbations, correlated representations, and minor model updates. Such variability undermines explanation reliability, as reliable explanations should remain consistent across equivalent input representations and small, performance-preserving model changes. We introduce the Explanation Reliability Index (ERI), a family of metrics that quantifies explanation stability under four reliability axioms: robustness to small input perturbations, consistency under feature redundancy, smoothness across model evolution, and resilience to mild distributional shifts. For each axiom, we derive formal guarantees, including Lipschitz-type bounds and temporal stability results. We further propose ERI-T, a dedicated measure of temporal reliability for sequential models, and introduce ERI-Bench, a benchmark designed to systematically stress-test explanation reliability across synthetic and real-world datasets. Experimental results reveal widespread reliability failures in popular explanation methods, showing that explanations can be unstable under realistic deployment conditions. By exposing and quantifying these instabilities, ERI enables principled assessment of explanation reliability and supports more trustworthy explainable AI (XAI) systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。