arXiv:2602.00063cs.LGcs.AI2026-02被引 2

模型不确定性会严重破坏因果解释的稳定性,需引入鲁棒性设计。

The Impact of Machine Learning Uncertainty on the Robustness of Counterfactual Explanations

  • 测试不同模型与解释算法在两类不确定下的表现
  • 模型准确率小幅下降即引发解释大幅变化
  • 金融、社科等领域尤其需要考虑不确定性

因果解释广泛用于解读机器学习预测,通过识别使模型决策改变的最小输入特征变动。然而,现有方法未在模型和数据不确定性变化下验证其稳健性,导致现实场景中解释可能不稳定或无效。本文研究常见模型与因果生成算法组合在似然不确定性和认知不确定性下的鲁棒性。基于合成与真实世界表格数据集的实验表明,因果解释对模型不确定性高度敏感。特别是,即使因噪声增加或数据有限导致模型准确率小幅下降,生成的因果解释在平均值和单个实例上也会出现显著差异。这一发现强调了在金融、社会科学等领域的解释方法中引入不确定性感知机制的必要性。

原文摘要 · Abstract (English)

Counterfactual explanations are widely used to interpret machine learning predictions by identifying minimal changes to input features that would alter a model's decision. However, most existing counterfactual methods have not been tested when model and data uncertainty change, resulting in explanations that may be unstable or invalid under real-world variability. In this work, we investigate the robustness of common combinations of machine learning models and counterfactual generation algorithms in the presence of both aleatoric and epistemic uncertainty. Through experiments on synthetic and real-world tabular datasets, we show that counterfactual explanations are highly sensitive to model uncertainty. In particular, we find that even small reductions in model accuracy - caused by increased noise or limited data - can lead to large variations in the generated counterfactuals on average and on individual instances. These findings underscore the need for uncertainty-aware explanation methods in domains such as finance and the social sciences.

因果解释不确定性模型鲁棒性可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。