arXiv:2411.15036cs.LGcs.SY2024-11被引 10

提出状态级安全约束的多智能体强化学习框架,确保每一步都安全。

Safe Multi-Agent Reinforcement Learning with Convergence to Generalized Nash Equilibrium

  • 基于可控不变集识别安全区域,实现每步状态的安全保障。
  • 算法收敛至广义纳什均衡,在高维系统中提升性能并减少违规。
  • 适合需全程安全的复杂多智能体应用,如自动驾驶协同控制。

多智能体强化学习(MARL)在合作任务中表现优异,但实际部署面临严峻安全挑战。现有安全MARL方法多基于约束马尔可夫决策过程(CMDP),仅对折扣累积成本施加约束,无法保证全程安全,且常忽略可行性问题——系统在某些约束集区域内必然违反状态约束,导致性能下降或违规增加。为此,本文提出一种新型理论框架,支持状态级(state-wise)约束,确保智能体访问的每个状态均满足安全要求。通过引入控制理论中的可控不变集(CIS)概念,以安全值函数刻画,设计多智能体方法识别CIS,并确保学习过程收敛至安全值函数上的纳什均衡。将CIS识别融入学习流程,提出多智能体对偶策略迭代算法,保证在状态约束合作马尔可夫博弈中收敛至广义纳什均衡,实现可行性和性能的最佳平衡。为适应复杂高维系统,进一步提出多智能体对偶演员-评论家(MADAC)算法,在深度强化学习范式下近似该迭代方案。在多个安全MARL基准测试中,MADAC持续优于现有方法,获得更高奖励并显著降低约束违规。

原文摘要 · Abstract (English)

Multi-agent reinforcement learning (MARL) has achieved notable success in cooperative tasks, demonstrating impressive performance and scalability. However, deploying MARL agents in real-world applications presents critical safety challenges. Current safe MARL algorithms are largely based on the constrained Markov decision process (CMDP) framework, which enforces constraints only on discounted cumulative costs and lacks an all-time safety assurance. Moreover, these methods often overlook the feasibility issue (the system will inevitably violate state constraints within certain regions of the constraint set), resulting in either suboptimal performance or increased constraint violations. To address these challenges, we propose a novel theoretical framework for safe MARL with $\textit{state-wise}$ constraints, where safety requirements are enforced at every state the agents visit. To resolve the feasibility issue, we leverage a control-theoretic notion of the feasible region, the controlled invariant set (CIS), characterized by the safety value function. We develop a multi-agent method for identifying CISs, ensuring convergence to a Nash equilibrium on the safety value function. By incorporating CIS identification into the learning process, we introduce a multi-agent dual policy iteration algorithm that guarantees convergence to a generalized Nash equilibrium in state-wise constrained cooperative Markov games, achieving an optimal balance between feasibility and performance. Furthermore, for practical deployment in complex high-dimensional systems, we propose $\textit{Multi-Agent Dual Actor-Critic}$ (MADAC), a safe MARL algorithm that approximates the proposed iteration scheme within the deep RL paradigm. Empirical evaluations on safe MARL benchmarks demonstrate that MADAC consistently outperforms existing methods, delivering much higher rewards while reducing constraint violations.

多智能体安全强化学习纳什均衡约束优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。