arXiv:2501.13727cs.MAcs.AI2025-01被引 4

提出可扩展安全多智能体强化学习框架,提升大规模系统安全性与效率

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System

  • 基于图结构设计分层消息传递网络,适应不同规模的局部观测与通信
  • 在局部观测下采用约束联合策略优化,显著提升系统安全性
  • 支持大规模智能体场景,性能优于最新方法,兼顾安全与效率

安全性和可扩展性是实际多智能体系统面临的两大挑战。现有仅依赖奖励塑造的多智能体强化学习(MARL)方法难以保证安全性,且因输出维度固定导致可扩展性差。为此,我们提出一种新框架——可扩展安全多智能体强化学习(SS-MARL),以提升MARL方法的安全性与可扩展性。利用多智能体系统的固有图结构,设计分层消息传递网络,实现对异构规模局部观测与通信信息的有效聚合;同时,在局部观测设定下提出约束联合策略优化方法,增强安全性。仿真实验表明,相比基线方法,SS-MARL在最优性与安全性之间取得更优平衡,且在大规模智能体场景下可扩展性显著优于最新方法。

原文摘要 · Abstract (English)

Safety and scalability are two critical challenges faced by practical Multi-Agent Systems (MAS). However, existing Multi-Agent Reinforcement Learning (MARL) algorithms that rely solely on reward shaping are ineffective in ensuring safety, and their scalability is rather limited due to the fixed-size network output. To address these issues, we propose a novel framework, Scalable Safe MARL (SS-MARL), to enhance the safety and scalability of MARL methods. Leveraging the inherent graph structure of MAS, we design a multi-layer message passing network to aggregate local observations and communications of varying sizes. Furthermore, we develop a constrained joint policy optimization method in the setting of local observation to improve safety. Simulation experiments demonstrate that SS-MARL achieves a better trade-off between optimality and safety compared to baselines, and its scalability significantly outperforms the latest methods in scenarios with a large number of agents.

多智能体强化学习安全可扩展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。