为自动驾驶车辆设计城市路径优化基准,推动强化学习应用
URB -- Urban Routing Benchmark for RL-equipped Connected Autonomous Vehicles
- 构建29个真实城市交通网络的统一评估环境
- 先进多智能体强化学习仍不如人类司机表现
- 适合研究大规模交通优化与强化学习的学者
联网自动驾驶车辆(CAVs)有望通过优化路径决策减少未来城市交通拥堵。与人类驾驶员不同,这些决策可基于集体、数据驱动的策略,借助机器学习算法实现。强化学习(RL)有助于开发此类协同路径策略,但缺乏标准化且真实的基准测试。为此,我们提出URB:面向强化学习赋能的联网自动驾驶车辆的城市路径优化基准。URB是一个综合性基准环境,统一了29个真实世界交通网络与现实需求模式下的评估。它包含预定义任务集、多智能体强化学习(MARL)算法实现、三种基线方法、领域特定性能指标及模块化配置方案。实验表明,尽管训练耗时且成本高昂,当前最先进的MARL算法极少优于人类表现。本文报告的实验结果开启了大规模城市路径优化中MARL的首个排行榜,揭示现有方法在扩展性上的困境,凸显该领域亟需突破。
原文摘要 · Abstract (English)
Connected Autonomous Vehicles (CAVs) promise to reduce congestion in future urban networks, potentially by optimizing their routing decisions. Unlike for human drivers, these decisions can be made with collective, data-driven policies, developed using machine learning algorithms. Reinforcement learning (RL) can facilitate the development of such collective routing strategies, yet standardized and realistic benchmarks are missing. To that end, we present URB: Urban Routing Benchmark for RL-equipped Connected Autonomous Vehicles. URB is a comprehensive benchmarking environment that unifies evaluation across 29 real-world traffic networks paired with realistic demand patterns. URB comes with a catalog of predefined tasks, multi-agent RL (MARL) algorithm implementations, three baseline methods, domain-specific performance metrics, and a modular configuration scheme. Our results show that, despite the lengthy and costly training, state-of-the-art MARL algorithms rarely outperformed humans. The experimental results reported in this paper initiate the first leaderboard for MARL in large-scale urban routing optimization. They reveal that current approaches struggle to scale, emphasizing the urgent need for advancements in this domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。