arXiv:2602.06521cs.CVcs.RO2026-02被引 24

让自动驾驶模型在潜空间中同时预测未来场景和规划动作,提升决策能力。

DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving

  • 在潜空间统一视觉-语言-动作与世界模型,共享潜在状态。
  • 在NAVSIMv1上达91.3的PDMS,在nuScenes上碰撞率仅0.16。
  • 适合做端到端自动驾驶系统研发与研究者参考。

端到端自动驾驶近年来受到关注,通过将视觉-语言-动作(VLA)与世界模型统一,以增强决策能力和未来想象。然而,现有方法因潜空间状态共享不足,难以有效融合未来场景演化与动作规划,限制了视觉想象对动作决策的影响。为此,我们提出DriveWorld-VLA,一种新框架,通过在表示层紧密集成VLA与世界模型,实现世界建模与规划在潜空间中的统一,使VLA规划器能直接受益于全局场景演化建模,并减少对密集标注监督的依赖。此外,该框架将世界模型的潜状态作为VLA规划器的核心决策状态,使规划器可评估候选动作对未来场景演化的影响。通过在潜空间内完成世界建模,DriveWorld-VLA支持可控、动作条件化的特征级想象,避免昂贵的像素级滚动推演。大量开环与闭环评估表明其有效性:在NAVSIMv1上取得91.3 PDMS,在NAVSIMv2上达86.8 EPDMS,nuScenes上3秒平均碰撞率为0.16。代码与模型将发布于https://github.com/liulin815/DriveWorld-VLA.git。

原文摘要 · Abstract (English)

End-to-end (E2E) autonomous driving has recently attracted increasing interest in unifying Vision-Language-Action (VLA) with World Models to enhance decision-making and forward-looking imagination. However, existing methods fail to effectively unify future scene evolution and action planning within a single architecture due to inadequate sharing of latent states, limiting the impact of visual imagination on action decisions. To address this limitation, we propose DriveWorld-VLA, a novel framework that unifies world modeling and planning within a latent space by tightly integrating VLA and world models at the representation level, which enables the VLA planner to benefit directly from holistic scene-evolution modeling and reducing reliance on dense annotated supervision. Additionally, DriveWorld-VLA incorporates the latent states of the world model as core decision-making states for the VLA planner, facilitating the planner to assess how candidate actions impact future scene evolution. By conducting world modeling entirely in the latent space, DriveWorld-VLA supports controllable, action-conditioned imagination at the feature level, avoiding expensive pixel-level rollouts. Extensive open-loop and closed-loop evaluations demonstrate the effectiveness of DriveWorld-VLA, which achieves state-of-the-art performance with 91.3 PDMS on NAVSIMv1, 86.8 EPDMS on NAVSIMv2, and 0.16 3-second average collision rate on nuScenes. Code and models will be released in https://github.com/liulin815/DriveWorld-VLA.git.

自动驾驶世界模型潜空间多模态规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。