arXiv:2512.19133cs.ROcs.CV2025-12AAAI被引 25

用强化微调优化视觉-几何世界模型,让自动驾驶更安全高效。

WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous Driving

论文配图:WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous Driving
图 1 · 摘自论文原文
  • 分层规划分解+局部迭代精修,引导表征学习面向决策优化
  • nuScenes上碰撞率降83%,仅0.05%;NavSim上性能媲美激光雷达方案
  • 适合追求高安全性、端到端自动驾驶系统的研究者与开发者

潜在世界模型通过时间自监督学习增强场景表征,为端到端自动驾驶提供无需感知标注的新范式。然而,以重建为导向的学习将感知与规划任务耦合,导致规划优化不佳。为此,我们提出WorldRFT——一种面向规划的潜在世界模型框架,通过分层规划分解与局部感知的交互式精修机制,并结合强化学习微调(RFT)提升关键策略的安全性。具体地,该框架融合视觉-几何基础模型以增强三维空间感知,利用分层规划任务分解引导表征优化,并通过局部感知的迭代精修生成面向规划的驾驶策略。此外,引入组相对策略优化(GRPO),结合轨迹高斯化与碰撞感知奖励,系统性提升安全性。WorldRFT在开环nuScenes与闭环NavSim基准上均达到最先进性能:在nuScenes上碰撞率降低83%(0.30% → 0.05%);在NavSim上仅使用相机输入,表现媲美基于激光雷达的SOTA方法DiffusionDrive(PDMS 87.8 vs. 88.1)。

原文摘要 · Abstract (English)

Latent World Models enhance scene representation through temporal self-supervised learning, presenting a perception annotation-free paradigm for end-to-end autonomous driving. However, the reconstruction-oriented representation learning tangles perception with planning tasks, leading to suboptimal optimization for planning. To address this challenge, we propose WorldRFT, a planning-oriented latent world model framework that aligns scene representation learning with planning via a hierarchical planning decomposition and local-aware interactive refinement mechanism, augmented by reinforcement learning fine-tuning (RFT) to enhance safety-critical policy performance. Specifically, WorldRFT integrates a vision-geometry foundation model to improve 3D spatial awareness, employs hierarchical planning task decomposition to guide representation optimization, and utilizes local-aware iterative refinement to derive a planning-oriented driving policy. Furthermore, we introduce Group Relative Policy Optimization (GRPO), which applies trajectory Gaussianization and collision-aware rewards to fine-tune the driving policy, yielding systematic improvements in safety. WorldRFT achieves state-of-the-art (SOTA) performance on both open-loop nuScenes and closed-loop NavSim benchmarks. On nuScenes, it reduces collision rates by 83% (0.30% -> 0.05%). On NavSim, using camera-only sensors input, it attains competitive performance with the LiDAR-based SOTA method DiffusionDrive (87.8 vs. 88.1 PDMS).

自动驾驶世界模型强化学习规划优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。