arXiv:2505.02293cs.ROcs.MA2025-05中稿 · publication at the…被引 15

用分层安全机制让多智能体强化学习更安全,避免碰撞冲突。

Resolving Conflicting Constraints in Multi-Agent Reinforcement Learning with Layered Safety

  • 结合MARL与安全滤波器,分层处理碰撞风险。
  • 硬件实验中冲突减少,且路径长度和时间接近基线。
  • 适合高密度无人机协同场景,强调安全与效率平衡。

多机器人导航中的防碰撞至关重要,但学习方法如多智能体强化学习(MARL)缺乏安全保证。传统控制方法在少数机器人交互时有效,但面对多智能体间约束冲突时失效。为此,本文提出分层安全MARL框架:首先,利用MARL学习减少多于两个智能体间的交互;其次,对进入相互作用距离的智能体对,优先处理最紧急的纠正需求;最后,通过专用安全滤波器提供战术修正动作。所有层级设计基于可达性分析与基于控制屏障值函数的过滤机制。我们在Crazyflie无人机硬件实验及高密度先进空中交通(AAM)场景中验证,结果表明该方法显著降低冲突,同时保持较高效率(路径更短、时间更少),优于无分层安全的基线方法。

原文摘要 · Abstract (English)

Preventing collisions in multi-robot navigation is crucial for deployment. This requirement hinders the use of learning-based approaches, such as multi-agent reinforcement learning (MARL), on their own due to their lack of safety guarantees. Traditional control methods, such as reachability and control barrier functions, can provide rigorous safety guarantees when interactions are limited only to a small number of robots. However, conflicts between the constraints faced by different agents pose a challenge to safe multi-agent coordination. To overcome this challenge, we propose a method that integrates multiple layers of safety by combining MARL with safety filters. First, MARL is used to learn strategies that minimize multiple agent interactions, where multiple indicates more than two. Particularly, we focus on interactions likely to result in conflicting constraints within the engagement distance. Next, for agents that enter the engagement distance, we prioritize pairs requiring the most urgent corrective actions. Finally, a dedicated safety filter provides tactical corrective actions to resolve these conflicts. Crucially, the design decisions for all layers of this framework are grounded in reachability analysis and a control barrier-value function-based filtering mechanism. We validate our Layered Safe MARL framework in 1) hardware experiments using Crazyflie drones and 2) high-density advanced aerial mobility (AAM) operation scenarios, where agents navigate to designated waypoints while avoiding collisions. The results show that our method significantly reduces conflict while maintaining safety without sacrificing much efficiency (i.e., shorter travel time and distance) compared to baselines that do not incorporate layered safety. The project website is available at https://dinamo-mit.github.io/Layered-Safe-MARL/

多智能体强化学习安全导航无人机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。