arXiv:2608.30956cs.CYcs.AI2026-08

Counterfactual解释的合理性受组织决策影响,不能忽视其背后的设计选择。

Taking the Whys Seriously: Limitations of Counterfactual Explanations in Justification and Recourse

论文配图:Taking the Whys Seriously: Limitations of Counterfactual Explanations in Justification and Recourse
图 1 · 摘自论文原文
  • 揭示了模型决策中测量、业务需求等上游选择对反事实解释的决定性影响
  • 实验证明这些设计选择对反事实结果的影响甚至超过生成方法本身
  • 提醒用户:反事实解释无法回答‘为何应如此决策’的根本问题

反事实解释(CEs)在可解释人工智能中被广泛用于展示输入特征变化如何影响模型输出,常用于调试模型、解释预测、提供决策理由及算法救济。本文探讨了在实际部署中使用反事实解释的规范正当性,发现其在说明与救济场景下需满足更严格标准。我们指出,简单应用反事实解释可能掩盖机器学习流程中诸多有争议的设计选择,使决策与反事实结果看似客观,实则反映组织的物质化设计与治理取向。通过四项实证实验,我们展示了对机器学习管道上游环节(如特征与标签的测量模型、业务要求、模型验证、成功度量标准)的干预会显著影响生成的反事实。结果表明,这些组织决策对反事实的影响程度不亚于甚至超过生成方法本身。因此,在提供解释与救济建议时必须考虑这些背景选择,凸显此类任务的相对性本质。作为理由或救济方案,反事实解释无法充分回应某些关键的‘为什么’问题,因其预设了既定决策不可质疑。

原文摘要 · Abstract (English)

Counterfactual explanations (CEs) are widely used in explainable artificial intelligence (AI) to show how a model's outputs would change if the input features were manipulated. This technique is used for a range of tasks such as debugging models, explaining predictions, justifying decisions, and providing algorithmic recourse. In this paper, we explore the normative legitimacy of employing counterfactuals in real-life model deployment settings. We discuss the different stakes involved in these different purposes for which CEs are commonly employed, and find stricter requirements for justification and recourse. In particular, we find that naive application of CEs for justification and recourse can lead to ignoring contestable choices made throughout the machine learning (ML) pipeline, thus obfuscating that decisions and counterfactuals for those decisions are also artifacts of an organization's materialized design and governance choices. We demonstrate this with four empirical experiments involving interventions at stages of the ML pipeline ``upstream" of the explanation itself, and show that these affect the generated counterfactuals. We find that an organization's choices on measurement models for feature and labels, business requirements, model validation, and the metric of model success have as much or more impact on the generated counterfactuals as the specifics of the generating method. Our findings underline the need to account for such choices upon providing justification and recourse, providing a stark reminder of the relational nature of these tasks. As putative justifications or recourse recommendations, CEs do not provide adequate answers to some important "why"-questions because they preclude consideration of whether the decision-maker ought to have acted differently.

可解释AI反事实解释决策公平模型伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。