arXiv:2510.19244cs.LG2025-10被引 1

让复杂智能体的决策变得可懂,提升信任度。

Interpret Policies in Deep Reinforcement Learning using SILVER with RL-Guided Labeling: A Model-level Approach to High-dimensional and Multi-action Environments

  • 用RL自身动作指导边界点识别,拓展解释能力至高维多动作场景
  • 在Atari上保持原有性能,同时显著提升人类对智能体行为的理解
  • 适合关注AI决策透明性与可信性的研究人员和工程师

深度强化学习(RL)表现优异但缺乏可解释性,限制了对策略行为的信任。现有SILVER框架(Li, Siddique, and Cao 2025)通过基于Shapley的回归解释策略,但仅适用于低维二元动作环境。本文提出SILVER结合RL引导标注的增强版本,通过将RL策略自身的动作输出融入边界点识别,拓展SILVER至多动作与高维环境。方法首先从图像观测中提取紧凑特征表示,进行基于SHAP的特征归因,再利用RL引导标注生成行为一致的边界数据集。随后训练决策树和基于回归的代理模型以揭示深度强化学习策略的决策结构。我们在两个Atari环境中使用三种深度强化学习算法评估该框架,并开展人机实验评估所推导解释策略的清晰度与可信度。结果表明,本方法在保持竞争性任务性能的同时,显著提升透明度与人类对智能体行为的理解。本工作推动可解释强化学习发展,将SILVER转化为适用于高维、多动作场景的可扩展且行为感知的解释框架。

原文摘要 · Abstract (English)

Deep reinforcement learning (RL) achieves remarkable performance but lacks interpretability, limiting trust in policy behavior. The existing SILVER framework (Li, Siddique, and Cao 2025) explains RL policy via Shapley-based regression but remains restricted to low-dimensional, binary-action domains. We propose SILVER with RL-guided labeling, an enhanced variant that extends SILVER to multi-action and high-dimensional environments by incorporating the RL policy's own action outputs into the boundary points identification. Our method first extracts compact feature representations from image observations, performs SHAP-based feature attribution, and then employs RL-guided labeling to generate behaviorally consistent boundary datasets. Surrogate models, such as decision trees and regression-based functions, are subsequently trained to interpret RL policy's decision structure. We evaluate the proposed framework on two Atari environments using three deep RL algorithms and conduct human-subject study to assess the clarity and trustworthiness of the derived interpretable policy. Results show that our approach maintains competitive task performance while substantially improving transparency and human understanding of agent behavior. This work advances explainable RL by transforming SILVER into a scalable and behavior-aware framework for interpreting deep RL agents in high-dimensional, multi-action settings.

可解释RL策略解释高维环境SHAP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。