让解释更可信:用不确定性保证生成可行动的反事实解释
CONFEX: Uncertainty-Aware Counterfactual Explanations with Conformal Guarantees
- 结合置信预测与整数规划,生成带不确定性的反事实解释
- 在多个数据集上证明解释更可靠且符合实际行为逻辑
- 适合需要高可信度解释的医疗、金融等关键领域应用
反事实解释(CFX)为模型预测提供人类可理解的理由,支持可操作的改进路径并提升可解释性。为确保可靠性,反事实解释应避开预测不确定性高的区域,否则可能导致误导或不适用。然而,现有方法常忽视不确定性,或缺乏将不确定性纳入的严谨机制及形式化保证。本文提出CONFEX,一种基于置信预测(CP)与混合整数线性规划(MILP)的新型不确定性感知反事实解释方法。CONFEX通过局部化置信预测程序,结合离线输入空间树分割,实现高效的MILP编码,从而生成具有局部覆盖率保证的解释,解决了传统反事实生成违反交换性的问题。实验表明,CONFEX在多个基准和指标上优于当前先进方法,其不确定性感知策略显著提升了解释的稳健性与合理性。
原文摘要 · Abstract (English)
Counterfactual explanations (CFXs) provide human-understandable justifications for model predictions, enabling actionable recourse and enhancing interpretability. To be reliable, CFXs must avoid regions of high predictive uncertainty, where explanations may be misleading or inapplicable. However, existing methods often neglect uncertainty or lack principled mechanisms for incorporating it with formal guarantees. We propose CONFEX, a novel method for generating uncertainty-aware counterfactual explanations using Conformal Prediction (CP) and Mixed-Integer Linear Programming (MILP). CONFEX explanations are designed to provide local coverage guarantees, addressing the issue that CFX generation violates exchangeability. To do so, we develop a novel localised CP procedure that enjoys an efficient MILP encoding by leveraging an offline tree-based partitioning of the input space. This way, CONFEX generates CFXs with rigorous guarantees on both predictive uncertainty and optimality. We evaluate CONFEX against state-of-the-art methods across diverse benchmarks and metrics, demonstrating that our uncertainty-aware approach yields robust and plausible explanations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。