无需人类示范,用自对弈强化学习让自动驾驶模型快速适应新城市。
Learning to Drive in New Cities Without Human Demonstrations
- 基于目标城市地图,通过自对弈多智能体强化学习训练驾驶策略。
- 在新城市任务成功率与轨迹真实性显著提升,无需任何本地人类数据。
- 适合需快速部署到新城市的自动驾驶团队,尤其适用于数据稀缺场景。
尽管自动驾驶车辆在特定运营区域内已实现可靠表现,但将其部署至新城市仍成本高昂且进展缓慢。主要瓶颈在于:当目标城市在道路几何、交通规则和交互模式上与训练数据差异较大时,需收集大量人类示范轨迹进行策略适配。本文提出无需数据的基于地图的自对弈自动驾驶(NOMAD),仅依赖目标城市地图与元信息,在仿真环境中完成驾驶策略迁移。通过简单奖励函数,NOMAD显著提升了新城市中的任务成功率与轨迹真实性,为数据密集型城市迁移方法提供了一种高效可扩展的替代方案。
原文摘要 · Abstract (English)
While autonomous vehicles have achieved reliable performance within specific operating regions, their deployment to new cities remains costly and slow. A key bottleneck is the need to collect many human demonstration trajectories when adapting driving policies to new cities that differ from those seen in training in terms of road geometry, traffic rules, and interaction patterns. In this paper, we show that self-play multi-agent reinforcement learning can adapt a driving policy to a substantially different target city using only the map and meta-information, without requiring any human demonstrations from that city. We introduce NO data Map-based self-play for Autonomous Driving (NOMAD), which enables policy adaptation in a simulator constructed based on the target-city map. Using a simple reward function, NOMAD substantially improves both task success rate and trajectory realism in target cities, demonstrating an effective and scalable alternative to data-intensive city-transfer methods. Project Page: https://nomaddrive.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。