用沙普利值将复杂强化学习策略转化为可解释的透明表示。
From Explainability to Interpretability: Interpretable Policies in Reinforcement Learning Via Model Explanation
- 通过沙普利值实现对策略的全局可解释性,超越局部解释。
- 在两个经典控制环境中保持原模型性能并生成更稳定的可解释策略。
- 适用于多种强化学习算法,适合需要信任决策过程的研究者。
深度强化学习在复杂领域表现出色,但深层神经网络策略的黑箱特性使得理解与信任其决策过程面临挑战。现有可解释强化学习方法仅提供局部洞察,难以实现全局理解,尤其在高风险应用中。为此,我们提出一种新型、模型无关的方法,利用沙普利值将复杂的深度强化学习策略转化为透明表示,弥合可解释性与可理解性之间的差距。该方法有两个关键贡献:一是基于沙普利值的策略解释新范式,突破局部解释局限;二是适用于离策略与在线策略算法的通用框架。我们在三种现有深度强化学习算法上进行评估,并在两个经典控制环境验证其有效性。结果表明,该方法不仅保持了原始模型性能,还生成了更稳定的可解释策略。
原文摘要 · Abstract (English)
Deep reinforcement learning (RL) has shown remarkable success in complex domains, however, the inherent black box nature of deep neural network policies raises significant challenges in understanding and trusting the decision-making processes. While existing explainable RL methods provide local insights, they fail to deliver a global understanding of the model, particularly in high-stakes applications. To overcome this limitation, we propose a novel model-agnostic approach that bridges the gap between explainability and interpretability by leveraging Shapley values to transform complex deep RL policies into transparent representations. The proposed approach offers two key contributions: a novel approach employing Shapley values to policy interpretation beyond local explanations and a general framework applicable to off-policy and on-policy algorithms. We evaluate our approach with three existing deep RL algorithms and validate its performance in two classic control environments. The results demonstrate that our approach not only preserves the original models' performance but also generates more stable interpretable policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。