arXiv:2604.01129cs.CV2026-04

让自动驾驶场景生成更可控,能模拟危险驾驶事故。

ReinDriveGen: Reinforcement Post-Training for Out-of-Distribution Driving Scene Generation

  • 用强化学习优化生成效果,提升异常场景表现。
  • 可自由编辑车辆轨迹,生成真实感强的驾驶视频。
  • 适合自动驾驶安全测试与仿真系统开发者。

我们提出ReinDriveGen,一种可完全控制动态驾驶场景的框架,支持用户自由编辑车辆轨迹,以模拟前车碰撞、车辆漂移、失控旋转、行人乱穿马路、自行车横穿车道等安全关键边缘情况。该方法从多帧激光雷达数据构建动态3D点云场景,引入车辆补全模块,从部分观测重建360°完整几何结构,并将编辑后的场景渲染为2D条件图像,引导视频扩散模型生成逼真的驾驶视频。由于这些编辑场景必然超出训练分布,我们进一步提出基于强化学习的后训练策略,结合成对偏好模型与成对奖励机制,在无真实标签监督下实现对分布外条件的鲁棒质量提升。大量实验表明,ReinDriveGen在编辑驾驶场景上优于现有方法,并在新视角合成任务中达到当前最佳性能。

原文摘要 · Abstract (English)

We present ReinDriveGen, a framework that enables full controllability over dynamic driving scenes, allowing users to freely edit actor trajectories to simulate safety-critical corner cases such as front-vehicle collisions, drifting cars, vehicles spinning out of control, pedestrians jaywalking, and cyclists cutting across lanes. Our approach constructs a dynamic 3D point cloud scene from multi-frame LiDAR data, introduces a vehicle completion module to reconstruct full 360° geometry from partial observations, and renders the edited scene into 2D condition images that guide a video diffusion model to synthesize realistic driving videos. Since such edited scenarios inevitably fall outside the training distribution, we further propose an RL-based post-training strategy with a pairwise preference model and a pairwise reward mechanism, enabling robust quality improvement under out-of-distribution conditions without ground-truth supervision. Extensive experiments demonstrate that ReinDriveGen outperforms existing approaches on edited driving scenarios and achieves state-of-the-art results on novel ego viewpoint synthesis.

自动驾驶视频生成强化学习场景仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。