arXiv:2601.16659cs.LGcs.AI2026-01

提出可证明鲁棒的反事实解释,应对模型更新带来的失效问题。

Provably Robust Bayesian Counterfactual Explanations under Model Changes

  • 基于贝叶斯框架,生成在模型变化下仍可靠的反事实解释。
  • 在多个数据集上验证,解释的预测置信度高且方差低。
  • 适合关注模型迭代时解释可信性的研究人员和开发者。

反事实解释(CEs)通过回答“如果……会怎样”来提供机器学习决策的可解释性。然而,在模型频繁更新的真实场景中,现有解释可能迅速失效或不可靠。本文提出概率安全反事实解释(PSCE),可生成δ-安全(保证高预测置信度)与ε-鲁棒(保证低预测方差)的解释。基于贝叶斯原理,PSCE在⟨δ, ε⟩集合内提供形式化概率保证。我们引入不确定性感知约束到优化框架中,并在多种数据集上进行实证验证。与最先进的贝叶斯方法相比,PSCE生成的解释不仅更合理、更具区分性,且在模型变更下具有可证明的鲁棒性。

原文摘要 · Abstract (English)

Counterfactual explanations (CEs) offer interpretable insights into machine learning predictions by answering ``what if?" questions. However, in real-world settings where models are frequently updated, existing counterfactual explanations can quickly become invalid or unreliable. In this paper, we introduce Probabilistically Safe CEs (PSCE), a method for generating counterfactual explanations that are $δ$-safe, to ensure high predictive confidence, and $ε$-robust to ensure low predictive variance. Based on Bayesian principles, PSCE provides formal probabilistic guarantees for CEs under model changes which are adhered to in what we refer to as the $\langle δ, ε\rangle$-set. Uncertainty-aware constraints are integrated into our optimization framework and we validate our method empirically across diverse datasets. We compare our approach against state-of-the-art Bayesian CE methods, where PSCE produces counterfactual explanations that are not only more plausible and discriminative, but also provably robust under model change.

反事实解释贝叶斯方法模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。