利用路侧单元的全局信息,让自动驾驶汽车更安全高效地完成匝道汇入。
Z-Merge: Multi-Agent Reinforcement Learning for On-Ramp Merging with Zone-Specific V2X Traffic Information
- 基于车路协同的多智能体强化学习框架,融合局部与全局信息决策。
- 在多种交通场景下,汇入成功率与通行效率显著提升,事故风险降低。
- 适合研究自动驾驶协同控制或智能交通系统的人参考。
匝道汇入是自动驾驶车辆在混合交通环境中面临的关键挑战。现有方法通常仅依赖局部或邻近车辆信息,采用变道或间隙创建策略,常导致安全性和通行效率不理想。本文提出一种基于车路协同(V2X)的多智能体强化学习框架,利用路侧单元(RSU)提供的区域特定全局信息,有效协调变道与车距调整策略。将汇入控制问题建模为多智能体部分可观马尔可夫决策过程(MA-POMDP),各智能体通过V2X通信融合本地与全局观测。为支持离散与连续控制动作,设计了混合动作空间,并采用参数化深度Q学习算法。结合SUMO交通仿真器与MOSAIC V2X仿真器的大量实验表明,该框架在多种交通场景中显著提升了汇入成功率、交通效率和道路安全性。
原文摘要 · Abstract (English)
Ramp merging is a critical and challenging task for autonomous vehicles (AVs), particularly in mixed traffic environments with human-driven vehicles (HVs). Existing approaches typically rely on either lane-changing or inter-vehicle gap creation strategies based solely on local or neighboring information, often leading to suboptimal performance in terms of safety and traffic efficiency. In this paper, we present a V2X (vehicle-to-everything communication)-assisted Multiagent Reinforcement Learning (MARL) framework for on-ramp merging that effectively coordinates the complex interplay between lane-changing and inter-vehicle gap adaptation strategies by utilizing zone-specific global information available from a roadside unit (RSU). The merging control problem is formulated as a Multiagent Partially Observable Markov Decision Process (MA-POMDP), where agents leverage both local and global observations through V2X communication. To support both discrete and continuous control decisions, we design a hybrid action space and adopt a parameterized deep Q-learning approach. Extensive simulations, integrating the SUMO traffic simulator and the MOSAIC V2X simulator, demonstrate that our framework significantly improves merging success rate, traffic efficiency, and road safety across diverse traffic scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。