让自动驾驶模型学会预判危险,显著提升安全性。
AD-R1: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving with Impartial World Models
- 用反事实数据生成真实碰撞场景,训练模型诚实预测风险。
- 在复杂模拟中使安全违规减少超50%,远超基线方法。
- 适合关注自动驾驶安全与强化学习落地的研究者。
端到端自动驾驶模型虽能直接从传感器数据学习复杂行为,但面临安全性和长尾事件处理难题。强化学习(RL)虽具潜力,却因世界模型固有的乐观偏差而难以突破。本文识别出该问题根源,并提出基于无偏世界模型的闭环强化学习框架。通过创新的反事实合成数据管道,系统生成大量可能发生的碰撞与离路事件,使模型从被动场景补全转向真实因果预测。该无偏模型作为内部评判者,在策略优化中引导智能体“梦”见动作后果。实验表明,该模型在新设立的危险预见基准上显著优于基线,且在挑战性仿真中安全违规降低超过50%,证明教会模型预想危险是实现真正安全智能自动驾驶的关键一步。
原文摘要 · Abstract (English)
End-to-end models for autonomous driving hold the promise of learning complex behaviors directly from sensor data, but face critical challenges in safety and handling long-tail events. Reinforcement Learning (RL) offers a promising path to overcome these limitations, yet its success in autonomous driving has been elusive. We identify a fundamental flaw hindering this progress: a deep seated optimistic bias in the world models used for RL. To address this, we introduce a framework for post-training policy refinement built around an Impartial World Model. Our primary contribution is to teach this model to be honest about danger. We achieve this with a novel data synthesis pipeline, Counterfactual Synthesis, which systematically generates a rich curriculum of plausible collisions and off-road events. This transforms the model from a passive scene completer into a veridical forecaster that remains faithful to the causal link between actions and outcomes. We then integrate this Impartial World Model into our closed-loop RL framework, where it serves as an internal critic. During refinement, the agent queries the critic to ``dream" of the outcomes for candidate actions. We demonstrate through extensive experiments, including on a new Risk Foreseeing Benchmark, that our model significantly outperforms baselines in predicting failures. Consequently, when used as a critic, it enables a substantial reduction in safety violations in challenging simulations, proving that teaching a model to dream of danger is a critical step towards building truly safe and intelligent autonomous agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。