用自动生成的挑战场景训练自动驾驶模型,减少事故。
Learning to Drive via Asymmetric Self-Play
- 教师生成学生无法解决的场景,学生学会应对这些难题。
- 在真实和罕见场景中碰撞率显著降低,优于现有方法。
- 适合需要提升泛化能力的自动驾驶研发人员。
大规模数据对学习真实且能力强的驾驶策略至关重要,但仅依赖真实数据存在局限:多数数据无意义,收集长尾场景成本高且不安全。本文提出非对称自对弈(asymmetric self-play),通过教师生成其能解决而学生无法解决的挑战性、可解且现实的合成场景,与学生协同训练。应用于交通仿真时,所学策略在常规和长尾场景中碰撞率显著下降。该策略还可零样本迁移至端到端自动驾驶训练数据生成,性能显著超越当前最优对抗方法或仅使用真实数据的方法。
原文摘要 · Abstract (English)
Large-scale data is crucial for learning realistic and capable driving policies. However, it can be impractical to rely on scaling datasets with real data alone. The majority of driving data is uninteresting, and deliberately collecting new long-tail scenarios is expensive and unsafe. We propose asymmetric self-play to scale beyond real data with additional challenging, solvable, and realistic synthetic scenarios. Our approach pairs a teacher that learns to generate scenarios it can solve but the student cannot, with a student that learns to solve them. When applied to traffic simulation, we learn realistic policies with significantly fewer collisions in both nominal and long-tail scenarios. Our policies further zero-shot transfer to generate training data for end-to-end autonomy, significantly outperforming state-of-the-art adversarial approaches, or using real data alone. For more information, visit https://waabi.ai/selfplay .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。