arXiv:2507.12599cs.AIcs.LG2025-07综述被引 8

梳理250篇论文,帮人看懂强化学习的决策逻辑

A Survey of Explainable Reinforcement Learning: Targets, Methods and Needs

  • 按解释目标和方式分类,构建清晰框架
  • 系统回顾超250篇文献,覆盖主流方法
  • 适合关注AI可解释性的研究者参考

近期人工智能模型的成功伴随其内部机制的不透明性,尤其源于深度神经网络的使用。为理解这些模型内部机制并解释其输出,一系列方法被提出,统称为可解释人工智能(XAI)。本文聚焦于XAI的一个子领域——可解释强化学习(XRL),旨在解释通过强化学习训练的智能体所采取的动作。我们基于两个问题——‘解释什么’与‘如何解释’——提出一个直观的分类体系:前者关注解释目标,后者关注解释方式。利用该分类体系,对超过250篇相关论文进行了综述。此外,我们还列举了若干与XRL密切相关的领域,认为应引起学术界重视。最后,指出了该领域当前存在的关键需求。

原文摘要 · Abstract (English)

The success of recent Artificial Intelligence (AI) models has been accompanied by the opacity of their internal mechanisms, due notably to the use of deep neural networks. In order to understand these internal mechanisms and explain the output of these AI models, a set of methods have been proposed, grouped under the domain of eXplainable AI (XAI). This paper focuses on a sub-domain of XAI, called eXplainable Reinforcement Learning (XRL), which aims to explain the actions of an agent that has learned by reinforcement learning. We propose an intuitive taxonomy based on two questions "What" and "How". The first question focuses on the target that the method explains, while the second relates to the way the explanation is provided. We use this taxonomy to provide a state-of-the-art review of over 250 papers. In addition, we present a set of domains close to XRL, which we believe should get attention from the community. Finally, we identify some needs for the field of XRL.

可解释性强化学习XAI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。