用谢林值统一解释强化学习的行为、结果和预测,让黑箱决策可懂可信。
A Theoretical Framework for Explaining Reinforcement Learning with Shapley Values
- 基于谢林值构建统一解释框架,量化环境特征对智能体决策的影响。
- 首次实现对行为、结果与预测的三重可解释性,且数学上严格合理。
- 适合需要高可信度的强化学习应用,如医疗、自动驾驶等安全场景。
强化学习智能体在复杂决策任务中可达到超人水平,但其行为往往难以理解与解释,限制了部署,尤其在对可解释性要求高的安全关键领域。本文识别出三个核心解释目标:行为、结果和预测,并提出一个统一的理论框架,通过分析智能体观察到的环境特征影响来解释这三类要素。该框架利用谢林值(Shapley values)计算特征贡献,其满足公平且一致的公理体系,确保信用分配合理。所提出的SVERL方法为强化学习智能体提供了一个整体、有意义且语义精确的解释框架,不仅可解释,且有数学依据,能揭示并修正先前解释中的概念缺陷。通过实例展示,SVERL能生成直观、有用的解释,揭示仅从行为本身无法察觉的决策逻辑。
原文摘要 · Abstract (English)
Reinforcement learning agents can achieve super-human performance in complex decision-making tasks, but their behaviour is often difficult to understand and explain. This lack of explanation limits deployment, especially in safety-critical settings where understanding and trust are essential. We identify three core explanatory targets that together provide a comprehensive view of reinforcement learning agents: behaviour, outcomes, and predictions. We develop a unified theoretical framework for explaining these three elements of reinforcement learning agents through the influence of individual features that the agent observes in its environment. We derive feature influences by using Shapley values, which collectively and uniquely satisfy a set of well-motivated axioms for fair and consistent credit assignment. The proposed approach, Shapley Values for Explaining Reinforcement Learning (SVERL), provides a single theoretical framework to comprehensively and meaningfully explain reinforcement learning agents. It yields explanations with precise semantics that are not only interpretable but also mathematically justified, enabling us to identify and correct conceptual issues in prior explanations. Through illustrative examples, we show how SVERL produces useful, intuitive explanations of agent behaviour, outcomes, and predictions, which are not apparent from observing agent behaviour alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。