arXiv:2603.14948cs.CV2026-03被引 10

统一视觉与运动表征,让自动驾驶模型既能生成未来场景又能实时规划路径。

Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation

  • 用轨迹词典约束视觉与运动一致性,生成与路径匹配的多模态未来场景。
  • 在nuScenes等数据集上,规划性能超越现有纯视觉方法,视频生成质量高。
  • 适合关注端到端自动驾驶系统整合与规划-生成协同优化的研究者。

端到端自动驾驶旨在从原始传感器输入中生成安全且合理的规划策略。驾驶世界模型通过预测驾驶场景的未来演变,展现出学习丰富表征的巨大潜力。然而,现有驾驶世界模型主要聚焦于视觉场景表征,运动表征未被显式设计为可共享、可继承的规划器输入,导致视觉生成优化与精确运动规划需求之间存在割裂。本文提出WorldDrive框架,通过统一视觉与运动表征,实现场景生成与实时规划的耦合。首先引入轨迹感知的驾驶世界模型,以轨迹词汇作为条件,强制视觉动态与运动意图一致,从而生成与特定轨迹相匹配的多样化、合理未来场景。将视觉和运动编码器迁移至下游多模态规划器,确保驾驶策略基于经过场景生成预优化的成熟表征。仅通过运动表征、视觉表征与自车状态的简单交互,即可生成高质量多模态轨迹。此外,为利用世界模型的前瞻性,提出未来感知奖励器,从冻结的世界模型中提取未来潜在表示,用于实时评估与选择最优轨迹。在NAVSIM、NAVSIM-v2和nuScenes基准上的大量实验表明,WorldDrive在纯视觉方法中达到领先的规划性能,同时保持高保真度的动作控制视频生成能力,有力证明了统一视觉与运动表征对鲁棒自动驾驶的有效性。

原文摘要 · Abstract (English)

End-to-end autonomous driving aims to generate safe and plausible planning policies from raw sensor input. Driving world models have shown great potential in learning rich representations by predicting the future evolution of a driving scene. However, existing driving world models primarily focus on visual scene representation, and motion representation is not explicitly designed to be planner-shared and inheritable, leaving a schism between the optimization of visual scene generation and the requirements of precise motion planning. We present WorldDrive, a holistic framework that couples scene generation and real-time planning via unifying vision and motion representation. We first introduce a Trajectory-aware Driving World Model, which conditions on a trajectory vocabulary to enforce consistency between visual dynamics and motion intentions, enabling the generation of diverse and plausible future scenes conditioned on a specific trajectory. We transfer the vision and motion encoders to a downstream Multi-modal Planner, ensuring the driving policy operates on mature representations pre-optimized by scene generation. A simple interaction between motion representation, visual representation, and ego status can generate high-quality, multi-modal trajectories. Furthermore, to exploit the world model's foresight, we propose a Future-aware Rewarder, which distills future latent representation from the frozen world model to evaluate and select optimal trajectories in real-time. Extensive experiments on the NAVSIM, NAVSIM-v2, and nuScenes benchmarks demonstrate that WorldDrive achieves leading planning performance among vision-only methods while maintaining high-fidelity action-controlled video generation capabilities, providing strong evidence for the effectiveness of unifying vision and motion representation for robust autonomous driving.

自动驾驶世界模型多模态规划端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。