用视频扩散模型重建真实驾驶场景,提升强化学习训练效果。
ReconDreamer-RL: Enhancing Reinforcement Learning via Diffusion-based Scene Reconstruction
- 结合视频扩散模型与运动学模型重建驾驶场景。
- 碰撞率降低5倍,优于模仿学习方法。
- 可自动生成罕见交通场景,适合自动驾驶研究者。
在闭环仿真中训练端到端自动驾驶模型的强化学习正日益受到关注。然而,多数仿真环境与真实世界存在显著差异,导致严重的仿真到现实(sim2real)差距。为缩小这一差距,部分方法利用场景重建技术生成逼真的模拟环境。尽管这提升了传感器仿真的真实性,但这些方法受限于训练数据分布,难以生成新轨迹或极端情况下的高质量传感器数据。为此,我们提出ReconDreamer-RL框架,将视频扩散先验引入场景重建以辅助强化学习,从而提升端到端自动驾驶训练效果。具体地,ReconDreamer-RL引入ReconSimulator,融合视频扩散先验进行外观建模,并结合运动学模型实现物理建模,从真实数据中重建驾驶场景,有效缩小闭环评估与强化学习中的sim2real差距。为覆盖更多极端场景,我们设计动态对抗代理(DAA),动态调整周边车辆相对于自车的轨迹,自主生成极端交通场景(如切入)。此外,提出表亲轨迹生成器(CTG)解决训练数据分布偏移问题,该问题常导致数据偏向简单直线行驶。实验表明,ReconDreamer-RL显著提升端到端自动驾驶训练效果,相较模仿学习方法碰撞率降低5倍。
原文摘要 · Abstract (English)
Reinforcement learning for training end-to-end autonomous driving models in closed-loop simulations is gaining growing attention. However, most simulation environments differ significantly from real-world conditions, creating a substantial simulation-to-reality (sim2real) gap. To bridge this gap, some approaches utilize scene reconstruction techniques to create photorealistic environments as a simulator. While this improves realistic sensor simulation, these methods are inherently constrained by the distribution of the training data, making it difficult to render high-quality sensor data for novel trajectories or corner case scenarios. Therefore, we propose ReconDreamer-RL, a framework designed to integrate video diffusion priors into scene reconstruction to aid reinforcement learning, thereby enhancing end-to-end autonomous driving training. Specifically, in ReconDreamer-RL, we introduce ReconSimulator, which combines the video diffusion prior for appearance modeling and incorporates a kinematic model for physical modeling, thereby reconstructing driving scenarios from real-world data. This narrows the sim2real gap for closed-loop evaluation and reinforcement learning. To cover more corner-case scenarios, we introduce the Dynamic Adversary Agent (DAA), which adjusts the trajectories of surrounding vehicles relative to the ego vehicle, autonomously generating corner-case traffic scenarios (e.g., cut-in). Finally, the Cousin Trajectory Generator (CTG) is proposed to address the issue of training data distribution, which is often biased toward simple straight-line movements. Experiments show that ReconDreamer-RL improves end-to-end autonomous driving training, outperforming imitation learning methods with a 5x reduction in the Collision Ratio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。