arXiv:2509.16500cs.CV2025-09NeurIPS被引 9

用几何反馈强化学习,让自动驾驶视频更真实可靠。

RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation

  • 用感知模型在潜在空间提供几何奖励,优化生成视频
  • 在nuScenes上使深度误差降低57%,3D检测准确率提升12.7%
  • 适合需要高精度几何结构的自动驾驶数据生成场景

合成数据对推动自动驾驶系统发展至关重要,但现有顶尖视频生成模型虽视觉逼真,仍存在细微几何失真,限制其在下游感知任务中的应用。我们识别并量化了这一关键问题,发现使用合成数据与真实数据在3D目标检测上存在显著性能差距。为此,提出基于几何反馈的强化学习方法RLGF,通过专用潜在空间感知模型提供奖励,精炼视频扩散模型。核心包括高效的潜在空间窗口优化技术,实现扩散过程中的精准反馈,以及分层几何奖励(HGR)系统,提供点-线-面对齐与场景占据一致性多级奖励。为量化失真,提出GeoScores。应用于DiVE模型在nuScenes数据集上,RLGF显著降低几何误差(如视角点误差下降21%,深度误差下降57%),3D目标检测平均精度(mAP)提升12.7%,大幅缩小与真实数据的差距。RLGF为自动驾驶开发提供了即插即用的高质量合成视频生成方案。

原文摘要 · Abstract (English)

Synthetic data is crucial for advancing autonomous driving (AD) systems, yet current state-of-the-art video generation models, despite their visual realism, suffer from subtle geometric distortions that limit their utility for downstream perception tasks. We identify and quantify this critical issue, demonstrating a significant performance gap in 3D object detection when using synthetic versus real data. To address this, we introduce Reinforcement Learning with Geometric Feedback (RLGF), RLGF uniquely refines video diffusion models by incorporating rewards from specialized latent-space AD perception models. Its core components include an efficient Latent-Space Windowing Optimization technique for targeted feedback during diffusion, and a Hierarchical Geometric Reward (HGR) system providing multi-level rewards for point-line-plane alignment, and scene occupancy coherence. To quantify these distortions, we propose GeoScores. Applied to models like DiVE on nuScenes, RLGF substantially reduces geometric errors (e.g., VP error by 21\%, Depth error by 57\%) and dramatically improves 3D object detection mAP by 12.7\%, narrowing the gap to real-data performance. RLGF offers a plug-and-play solution for generating geometrically sound and reliable synthetic videos for AD development.

自动驾驶视频生成强化学习几何对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。