arXiv:2603.27287cs.ROcs.CV2026-03被引 6

将环境预测与路径规划交替进行,提升自动驾驶决策准确性

Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving

  • 预测与规划交替进行,形成闭环交互
  • 在NAVSIM上实现高保真预测与竞争力规划性能
  • 融入单目深度信息,增强长时序预测能力

自动驾驶需要推理环境演变并据此规划行动。现有基于世界模型的方法通常先预测未来场景再进行规划,导致开环想象可能偏离实际决策过程。本文提出Uni-World VLA,一种统一的视觉-语言-动作(VLA)模型,将未来帧预测与轨迹规划紧密交织。模型不预先生成完整世界演化序列,而是逐步交替预测未来帧与自车动作,使规划决策持续依赖于所设想的未来观测。这种交替生成构建了世界建模与控制之间的闭环交互,增强了动态交通场景下的自适应决策能力。此外,模型引入单目深度信息,为世界建模提供更强几何线索,提升了长时序场景预测性能。在NAVSIM基准上的实验表明,该方法在保持高保真未来帧预测的同时,实现了具有竞争力的闭环规划表现。结果证明,紧密耦合世界预测与规划是可扩展VLA自动驾驶系统的一个有前景方向。

原文摘要 · Abstract (English)

Autonomous driving requires reasoning about how the environment evolves and planning actions accordingly. Existing world-model-based approaches typically predict future scenes first and plan afterwards, resulting in open-loop imagination that may drift from the actual decision process. In this paper, we present Uni-World VLA, a unified vision-language-action (VLA) model that tightly interleaves future frame prediction and trajectory planning. Instead of generating a full world rollout before planning, our model alternates between predicting future frames and ego actions step by step, allowing planning decisions to be continuously conditioned on the imagined future observations. This interleaved generation forms a closed-loop interaction between world modeling and control, enabling more adaptive decision-making in dynamic traffic scenarios. In addition, we incorporate monocular depth information into frames to provide stronger geometric cues for world modeling, improving long-horizon scene prediction. Experiments on the NAVSIM benchmark show that our approach achieves competitive closed-loop planning performance while producing high-fidelity future frame predictions. These results demonstrate that tightly coupling world prediction and planning is a promising direction for scalable VLA driving systems.

自动驾驶世界模型规划闭环

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。