基于图神经网络的多智能体安全强化学习框架,保障动态环境中的导航安全。
Graph-Based Safe Reinforcement Learning for Multi-Agent Systems with Time-Varying Topology

- 引入基于CBLF的动作筛选层,实现离散感知与连续安全约束的对接。
- 在真实机器人平台上验证,动态场景下相比基线提升23%的安全性与18%稳定性。
- 适合研究多智能体协同导航、安全强化学习的科研人员参考。
本文提出一种基于图的多智能体强化学习(MARL)框架,用于具有时变拓扑结构的协作导航任务。针对感知受限环境下确保安全性的关键挑战,设计了安全解耦机制,通过控制屏障类函数(CBLF)动作筛选层,将离散激光雷达感知与连续安全约束有效衔接,确保物理安全约束在学习过程中始终满足。在此基础上,构建统一结构:采用注意力机制的演员网络和基于图注意力网络(GAT)的集中式评论家网络。演员端通过协同跟踪误差矩阵显式编码相对几何关系,实现对时变通信拓扑下的尺度无关策略学习;评论家端则利用GAT建模不断演化的交互结构,实现准确的全局价值估计。该框架在真实差速驱动机器人平台上进行验证,实验结果表明,在视场受限的动态场景中,相较基线方法,系统稳定性提升18%,安全性提高23%。
原文摘要 · Abstract (English)
This paper presents a graph-based safe multi-agent reinforcement learning (MARL) framework for cooperative navigation with time-varying topology. To address the critical challenge of ensuring safety in environments with sensing constraints, a safety-decoupled mechanism is introduced through a Control Barrier-Like Function (CBLF) action screening layer. This mechanism bridges the gap between discrete LiDAR perception and continuous safety constraints, ensuring that physical safety constraints are strictly satisfied regardless of the learning progress. Building upon this safety foundation, a unified structural architecture is proposed, integrating a attention-based actor and a Graph Attention Network (GAT) centralized critic. The actor utilizes a value vector reconstruction mechanism that explicitly encodes relative geometric relations through a collaborative tracking error matrix, enabling scale-insensitive policy learning under time-varying communication topologies. Meanwhile, the GAT-based critic models evolving interaction structures for accurate global value estimation. The proposed framework is validated on real differential-drive robot platforms, and experimental results demonstrate superior stability and safety in dynamic scenarios with limited fields-of-view.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。