用行动序列定义反事实解释,让AI的决策过程更可理解。
Counterfactual Explanations as Plans
- 将反事实解释建模为一系列可执行的动作序列。
- 在部分真相、弱化真相和错误信念下均能有效工作。
- 适合需要解释复杂决策链的AI系统设计者。
近年来,人工智能可解释性受到广泛关注,尤其针对黑箱机器学习模型。正如规划领域所指出的,当任务不是单次决策,而是依赖观察的连续动作序列时,需要更丰富的解释方式。本文基于动作序列形式化了“反事实解释”,并自然引出模型一致性修正机制:用户可纠正代理的模型,或建议其行动计划。为此,需区分“事实”与“已知”,我们采用情境演算的模态片段来形式化这些直觉。考虑多种场景:代理知晓部分真相、弱化真相或持有错误信念,证明该框架可轻松推广至各类情形。
原文摘要 · Abstract (English)
There has been considerable recent interest in explainability in AI, especially with black-box machine learning models. As correctly observed by the planning community, when the application at hand is not a single-shot decision or prediction, but a sequence of actions that depend on observations, a richer notion of explanations are desirable. In this paper, we look to provide a formal account of ``counterfactual explanations," based in terms of action sequences. We then show that this naturally leads to an account of model reconciliation, which might take the form of the user correcting the agent's model, or suggesting actions to the agent's plan. For this, we will need to articulate what is true versus what is known, and we appeal to a modal fragment of the situation calculus to formalise these intuitions. We consider various settings: the agent knowing partial truths, weakened truths and having false beliefs, and show that our definitions easily generalize to these different settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。