arXiv:2504.15425cs.ROcs.AI2025-04中稿 · Robotics: Science …被引 10

提出分布式安全强化学习算法,让多机器人稳定协作完成任务且零违规。

Solving Multi-Agent Safe Optimal Control with Distributed Epigraph Form MARL

  • 将约束优化转为等价的上图形式,提升训练稳定性。
  • 在8个仿真任务中实现零约束违规,训练更稳定。
  • 适合需要高安全性协作的多机器人系统,如无人机编队。

多机器人系统任务常需协作完成团队目标并保证安全。该问题通常建模为约束马尔可夫决策过程(CMDP),目标是最小化全局代价并使约束违反的均值低于用户设定阈值。受真实机器人应用启发,本文将安全定义为零约束违反。尽管已有多种安全多智能体强化学习(MARL)算法用于求解CMDP,但这些方法在此设置下训练不稳定。为此,本文采用约束优化的上图形式以提升训练稳定性,并证明中心化上图问题可通过各智能体分布式求解。由此提出一种新的集中训练、分布式执行的MARL算法——Def-MARL。在两个不同模拟器上的8个任务的仿真实验表明,Def-MARL性能最优,满足安全约束且训练稳定。在Crazyflie四旋翼飞行器上的真实硬件实验进一步验证了其在复杂协作任务中安全协调多智能体的能力,优于其他方法。

原文摘要 · Abstract (English)

Tasks for multi-robot systems often require the robots to collaborate and complete a team goal while maintaining safety. This problem is usually formalized as a constrained Markov decision process (CMDP), which targets minimizing a global cost and bringing the mean of constraint violation below a user-defined threshold. Inspired by real-world robotic applications, we define safety as zero constraint violation. While many safe multi-agent reinforcement learning (MARL) algorithms have been proposed to solve CMDPs, these algorithms suffer from unstable training in this setting. To tackle this, we use the epigraph form for constrained optimization to improve training stability and prove that the centralized epigraph form problem can be solved in a distributed fashion by each agent. This results in a novel centralized training distributed execution MARL algorithm named Def-MARL. Simulation experiments on 8 different tasks across 2 different simulators show that Def-MARL achieves the best overall performance, satisfies safety constraints, and maintains stable training. Real-world hardware experiments on Crazyflie quadcopters demonstrate the ability of Def-MARL to safely coordinate agents to complete complex collaborative tasks compared to other methods.

多智能体安全强化学习分布式优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。