arXiv:2502.06869cs.LGcs.AI2025-02综述被引 21

解析深度强化学习的黑箱问题,提升决策透明度与可信度。

A Survey on Explainable Deep Reinforcement Learning

  • 从特征、状态、数据到模型层面,提供多维度解释方法。
  • 评估框架涵盖定性与定量指标,支持策略优化与安全增强。
  • 适合关注AI可解释性、安全部署的研究者与工程师。

深度强化学习(DRL)在各类序列决策任务中取得显著进展,但其依赖黑箱神经网络架构,限制了可解释性、信任度和高风险场景下的应用。可解释深度强化学习(XRL)通过特征级、状态级、数据集级和模型级的解释技术,提升系统透明性。本综述全面梳理了XRL方法,评估其定性与定量评估框架,并探讨其在策略优化、对抗鲁棒性与安全性中的作用。此外,研究还分析了强化学习与大语言模型(LLMs)的融合,特别是基于人类反馈的强化学习(RLHF),用于对齐人工智能与人类偏好。最后,总结开放挑战与未来方向,推动可解释、可靠且可问责的DRL系统发展。

原文摘要 · Abstract (English)

Deep Reinforcement Learning (DRL) has achieved remarkable success in sequential decision-making tasks across diverse domains, yet its reliance on black-box neural architectures hinders interpretability, trust, and deployment in high-stakes applications. Explainable Deep Reinforcement Learning (XRL) addresses these challenges by enhancing transparency through feature-level, state-level, dataset-level, and model-level explanation techniques. This survey provides a comprehensive review of XRL methods, evaluates their qualitative and quantitative assessment frameworks, and explores their role in policy refinement, adversarial robustness, and security. Additionally, we examine the integration of reinforcement learning with Large Language Models (LLMs), particularly through Reinforcement Learning from Human Feedback (RLHF), which optimizes AI alignment with human preferences. We conclude by highlighting open research challenges and future directions to advance the development of interpretable, reliable, and accountable DRL systems.

可解释性强化学习AI安全大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。