提出可解释的反事实解释排序方法,让AI决策更透明可信。
Ranking Counterfactual Explanations
- 定义反事实解释并建立严格排序机制,超越简单最小化条件。
- 12个真实数据集实验表明多数案例仅有一个最优解释。
- 最优解释更具代表性,能覆盖更广的样本范围,适合可信AI研究者。
AI决策结果常难以被用户理解。解释需回答两个问题:为何是这个结果(事实性)与为何不是另一个(反事实性)。尽管已有大量工作形式化事实性解释,但对反事实解释的系统研究仍不足。本文提出反事实解释的形式化定义,证明其满足的性质,并分析其与事实性解释的关系。由于同一案例通常存在多个反事实解释,我们引入一种严谨的排序方法,以识别最优解释,超越简单的最小化条件。在12个真实数据集上的实验表明,大多数情况下仅存在一个最优反事实解释。通过三个指标验证,所选最优解释具有更高的代表性,能解释更广泛的样本元素。该结果凸显了本方法在识别更稳健、全面反事实解释方面的有效性。
原文摘要 · Abstract (English)
AI-driven outcomes can be challenging for end-users to understand. Explanations can address two key questions: "Why this outcome?" (factual) and "Why not another?" (counterfactual). While substantial efforts have been made to formalize factual explanations, a precise and comprehensive study of counterfactual explanations is still lacking. This paper proposes a formal definition of counterfactual explanations, proving some properties they satisfy, and examining the relationship with factual explanations. Given that multiple counterfactual explanations generally exist for a specific case, we also introduce a rigorous method to rank these counterfactual explanations, going beyond a simple minimality condition, and to identify the optimal ones. Our experiments with 12 real-world datasets highlight that, in most cases, a single optimal counterfactual explanation emerges. We also demonstrate, via three metrics, that the selected optimal explanation exhibits higher representativeness and can explain a broader range of elements than a random minimal counterfactual. This result highlights the effectiveness of our approach in identifying more robust and comprehensive counterfactual explanations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。