arXiv:2604.15308cs.CV2026-04被引 4

用生成器-判别器框架提升自动驾驶规划的稳定性与安全性

RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework

论文配图:RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework
图 1 · 摘自论文原文
  • 生成器产多样轨迹,判别器用强化学习重排优劣
  • 碰撞率降低56%,实测城市路况更安全平滑
  • 适合做高阶自动驾驶规划系统的研究与工程落地

高等级自动驾驶需要具备建模多模态未来不确定性并保持闭环交互鲁棒性的运动规划能力。尽管基于扩散模型的规划器能有效捕捉复杂轨迹分布,但纯模仿学习训练时常出现随机不稳和缺乏纠正性负反馈的问题。为此,我们提出RAD-2,一种统一的生成器-判别器框架用于闭环规划:扩散生成器产出多样化轨迹候选,而强化学习优化的判别器根据长期驾驶质量对候选进行重排序。该解耦设计避免直接在高维轨迹空间施加稀疏标量奖励,从而提升优化稳定性。为进一步增强强化学习,我们引入时间一致的组相对策略优化,利用时间一致性缓解信用分配难题;同时提出在线策略生成器优化,将闭环反馈转化为结构化的纵向优化信号,逐步引导生成器逼近高奖励轨迹流形。为支持高效大规模训练,我们设计了BEV-Warp,一种高吞吐仿真环境,通过空间扭曲在鸟瞰图特征空间中直接执行闭环评估。RAD-2相比强基线扩散规划器碰撞率降低56%。真实世界部署进一步验证其在复杂城市交通中提升了感知安全性和驾驶平顺性。

原文摘要 · Abstract (English)

High-level autonomous driving requires motion planners capable of modeling multimodal future uncertainties while remaining robust in closed-loop interactions. Although diffusion-based planners are effective at modeling complex trajectory distributions, they often suffer from stochastic instabilities and the lack of corrective negative feedback when trained purely with imitation learning. To address these issues, we propose RAD-2, a unified generator-discriminator framework for closed-loop planning. Specifically, a diffusion-based generator is used to produce diverse trajectory candidates, while an RL-optimized discriminator reranks these candidates according to their long-term driving quality. This decoupled design avoids directly applying sparse scalar rewards to the full high-dimensional trajectory space, thereby improving optimization stability. To further enhance reinforcement learning, we introduce Temporally Consistent Group Relative Policy Optimization, which exploits temporal coherence to alleviate the credit assignment problem. In addition, we propose On-policy Generator Optimization, which converts closed-loop feedback into structured longitudinal optimization signals and progressively shifts the generator toward high-reward trajectory manifolds. To support efficient large-scale training, we introduce BEV-Warp, a high-throughput simulation environment that performs closed-loop evaluation directly in Bird's-Eye View feature space via spatial warping. RAD-2 reduces the collision rate by 56% compared with strong diffusion-based planners. Real-world deployment further demonstrates improved perceived safety and driving smoothness in complex urban traffic.

自动驾驶强化学习扩散模型规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。