让自动驾驶车在真实道路安全协作,靠通信+鲁棒算法实现仿真到硬件的零样本迁移。
Robust and Safe Multi-Agent Reinforcement Learning with Communication for Autonomous Vehicles: From Simulation to Hardware
- 设计通信增强的鲁棒安全多智能体强化学习框架
- 1/10比例实车实验验证跨场景安全协同能力
- 结合控制屏障函数保障个体安全,适合高阶自动驾驶研究
深度多智能体强化学习(MARL)在多机器人仿真中表现优异。对于自动驾驶车辆,车对车(V2V)通信技术为提升系统安全性提供了新机遇。然而,将仿真训练的MARL策略直接迁移到动态真实系统仍面临挑战,且如何利用通信与共享信息实现高效MARL在真实硬件上的演示仍有限。这一难题源于仿真与物理状态差异、系统状态与模型不确定性、实用共享信息设计以及仿真和硬件中均需保证安全性的需求。本文提出新型鲁棒安全多智能体强化学习框架RSR-RSMARL,支持真实-仿真-真实(RSR)策略适配,具备智能体间通信能力,并完成仿真与硬件双重验证。该框架考虑真实系统复杂性,采用状态(含智能体间共享状态信息)与动作表示进行MARL建模,通过鲁棒MARL算法训练策略以实现对硬件的零样本迁移。每个智能体配备基于控制屏障函数(CBFs)的安全盾模块,提供个体安全保证。在1/10比例自动驾驶车辆与V2V通信的实验中,结果表明该框架可显著提升驾驶安全性和多配置下的协同性能。研究强调了联合设计鲁棒策略表示与模块化安全架构对实现可扩展、泛化性强的多智能体自主系统真-仿-真迁移的重要性。
原文摘要 · Abstract (English)
Deep multi-agent reinforcement learning (MARL) has been demonstrated effectively in simulations for multi-robot problems. For autonomous vehicles, the development of vehicle-to-vehicle (V2V) communication technologies provide opportunities to further enhance system safety. However, zero-shot transfer of simulator-trained MARL policies to dynamic hardware systems remains challenging, and how to leverage communication and shared information for MARL has limited demonstrations on hardware. This problem is challenged by discrepancies between simulated and physical states, system state and model uncertainties, practical shared information design, and the need for safety guarantees in both simulation and hardware. This paper designs RSR-RSMARL, a novel Robust and Safe MARL framework that supports Real-Sim-Real (RSR) policy adaptation for multi-agent systems with communication among agents, with both simulation and hardware demonstrations. RSR-RSMARL leverages state (includes shared state information among agents) and action representations considering real system complexities for MARL formulation. The MARL policy is trained with robust MARL algorithm to enable zero-shot transfer to hardware considering the sim-to-real gap. A safety shield module using Control Barrier Functions (CBFs) provides safety guarantee for each individual agent. Experimental results on 1/10th-scale autonomous vehicles with V2V communication demonstrate the ability of RSR-RSMARL framework to enhance driving safety and coordination across multiple configurations. These findings emphasize the importance of jointly designing robust policy representations and modular safety architectures to enable scalable, generalizable RSR transfer in multi-agent autonomy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。