arXiv:2505.17342cs.LG2025-05综述被引 28

系统梳理安全强化学习的理论与算法,聚焦单/多智能体安全约束。

A Survey of Safe Reinforcement Learning and Constrained MDPs: A Technical Survey on Single-Agent and Multi-Agent Safety

  • 基于约束马尔可夫决策过程建模安全约束,统一理论框架。
  • 总结单智能体安全策略梯度与探索方法,涵盖协同与竞争场景的多智能体进展。
  • 提出5个开放问题,重点推进多智能体安全研究。

安全强化学习(SafeRL)是强化学习中专门处理智能体学习与部署过程中安全约束的子领域。本综述从数学严谨性出发,基于约束马尔可夫决策过程(CMDPs)及其在多智能体安全强化学习(SafeMARL)中的扩展,系统梳理了相关理论基础,包括定义、约束优化技术与基本定理。随后,总结了单智能体安全强化学习的前沿算法,如具有安全保证的策略梯度方法和安全探索策略,并介绍近期在协作与竞争场景下的多智能体安全强化学习进展。此外,本文提出了五个开放研究问题,其中三个聚焦于SafeMARL,每个问题均包含动机、关键挑战及前期工作关联。该综述旨在为关注安全强化学习与多智能体安全的研究者提供技术指南,明确核心概念、方法及未来研究方向。

原文摘要 · Abstract (English)

Safe Reinforcement Learning (SafeRL) is the subfield of reinforcement learning that explicitly deals with safety constraints during the learning and deployment of agents. This survey provides a mathematically rigorous overview of SafeRL formulations based on Constrained Markov Decision Processes (CMDPs) and extensions to Multi-Agent Safe RL (SafeMARL). We review theoretical foundations of CMDPs, covering definitions, constrained optimization techniques, and fundamental theorems. We then summarize state-of-the-art algorithms in SafeRL for single agents, including policy gradient methods with safety guarantees and safe exploration strategies, as well as recent advances in SafeMARL for cooperative and competitive settings. Additionally, we propose five open research problems to advance the field, with three focusing on SafeMARL. Each problem is described with motivation, key challenges, and related prior work. This survey is intended as a technical guide for researchers interested in SafeRL and SafeMARL, highlighting key concepts, methods, and open future research directions.

安全强化学习约束MDP多智能体综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。