arXiv:2507.14850cs.LGcs.AI2025-07NeurIPS被引 7

用分层强化学习+安全函数,让多个智能体在复杂道路中零事故协同通行。

Hierarchical Multi-Agent Reinforcement Learning with Control Barrier Functions for Safety-Critical Autonomous Systems

  • 分层设计:高层学协作技能,底层用安全函数保障单个智能体行为安全
  • 在多车冲突道路场景中实现接近100%安全率(误差小于5%)
  • 适合自动驾驶、无人机群等高安全性要求的多智能体系统

我们研究多智能体安全关键自主系统中的安全策略学习问题。在此类系统中,每个智能体必须始终满足安全要求,同时与其他智能体协作完成任务。为此,我们提出一种基于控制屏障函数(CBFs)的安全分层多智能体强化学习(HMARL)方法。该方法将整体强化学习问题分解为两层:高层学习所有智能体的联合协作策略,底层在高层策略条件下学习各智能体的安全执行策略。具体地,我们提出一种基于技能的HMARL-CBF算法,高层学习跨智能体技能的联合策略,底层则利用CBFs学习安全执行技能的策略。我们在大量智能体需在冲突道路网络中安全导航的挑战性场景中验证了该方法。相比现有最先进方法,本方法显著提升安全性,在所有环境中均达到近似完美的安全成功率(误差小于5%),同时性能全面优于对比方法。

原文摘要 · Abstract (English)

We address the problem of safe policy learning in multi-agent safety-critical autonomous systems. In such systems, it is necessary for each agent to meet the safety requirements at all times while also cooperating with other agents to accomplish the task. Toward this end, we propose a safe Hierarchical Multi-Agent Reinforcement Learning (HMARL) approach based on Control Barrier Functions (CBFs). Our proposed hierarchical approach decomposes the overall reinforcement learning problem into two levels learning joint cooperative behavior at the higher level and learning safe individual behavior at the lower or agent level conditioned on the high-level policy. Specifically, we propose a skill-based HMARL-CBF algorithm in which the higher level problem involves learning a joint policy over the skills for all the agents and the lower-level problem involves learning policies to execute the skills safely with CBFs. We validate our approach on challenging environment scenarios whereby a large number of agents have to safely navigate through conflicting road networks. Compared with existing state of the art methods, our approach significantly improves the safety achieving near perfect (within 5%) success/safety rate while also improving performance across all the environments.

多智能体强化学习安全控制自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。