arXiv:2502.03349cs.LGcs.AI2025-02ICML被引 72

通过自对弈训练,自动驾驶系统在仿真中实现了超大规模的自然驾驶行为。

Robust Autonomy Emerges from Self-Play

  • 利用自对弈机制,在仿真中让智能体相互博弈进化驾驶策略。
  • 模拟行驶16亿公里,平均17.5年才发生一次事故,性能超越现有最佳方法。
  • 无需人类驾驶数据,即可在真实交通环境中稳定运行,适合高安全性自动驾驶研究。

自对弈已推动双人及多人游戏的突破性进展。本文首次证明,自对弈在另一领域同样高效:在前所未有的规模下,仅通过自对弈训练,仿真中涌现出稳健且自然的自动驾驶行为——总行驶里程达16亿公里。这得益于Gigaflow,一个批处理仿真器,可在单个8卡节点上每小时合成并训练42年人类主观驾驶经验。所获策略在三个独立自动驾驶基准测试中达到当前最优表现;在真实世界录制场景中,面对人类车辆时仍优于以往最佳模型,且训练过程中从未接触过人类数据。该策略在与人类参考对比时表现出高度真实性,并实现前所未有的鲁棒性,仿真中平均连续行驶17.5年才发生一次事故。

原文摘要 · Abstract (English)

Self-play has powered breakthroughs in two-player and multi-player games. Here we show that self-play is a surprisingly effective strategy in another domain. We show that robust and naturalistic driving emerges entirely from self-play in simulation at unprecedented scale -- 1.6~billion~km of driving. This is enabled by Gigaflow, a batched simulator that can synthesize and train on 42 years of subjective driving experience per hour on a single 8-GPU node. The resulting policy achieves state-of-the-art performance on three independent autonomous driving benchmarks. The policy outperforms the prior state of the art when tested on recorded real-world scenarios, amidst human drivers, without ever seeing human data during training. The policy is realistic when assessed against human references and achieves unprecedented robustness, averaging 17.5 years of continuous driving between incidents in simulation.

自动驾驶自对弈仿真训练强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。