自博弈训练自动驾驶策略时,哪些交通规则会自然出现?
What Emerges and What Breaks in Self-Play Driving

- 用Transformer在真实城市高清地图上进行自博弈训练
- 发现模型在红绿灯处存在奖励漏洞,对停车标志无响应
- 验证奖励设计能实现预期的多样化驾驶行为
通过纯自博弈方式训练自动驾驶策略近期取得显著进展。受Gigaflow和Puffer-Drive启发,我们采用Transformer架构并在真实城市的高精度地图上进行训练,目标是最终部署。在CARLA和Waymax基准测试中,我们的策略表现不及Gigaflow,分析发现失败模式包括红绿灯处的奖励滥用以及缺乏在停车标志前停止的激励。进一步研究自博弈过程中自然涌现的交通规则及其与人类驾驶的契合度,并证实奖励调节可实现预期的行为多样性。训练策略演示见https://laursisask-ut.github.io/eccvdemo。
原文摘要 · Abstract (English)
Training autonomous driving policies through pure self-play has recently shown promising results. Following Gigaflow and Puffer- Drive, we train driving policies in a similar self-play fashion, but extend the models from MLPs to Transformers and train on the high-definition map of a real city, where we ultimately aim to deploy them. On the CARLA and Waymax benchmarks, our policies fall short of Gigaflow, and we trace the gap to specific failure modes, including reward hacking at traffic lights and a missing incentive to stop at stop signs. We further analyze which traffic rules emerge from self-play and how closely they match human driving, and we confirm that reward conditioning yields the intended diversity of driving behaviors. A demonstration of a trained policy is available at https://laursisask-ut.github.io/eccvdemo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。