将视觉语言模型与世界模型结合,提升自动驾驶的场景理解与动态预测能力。
WorldVLM: Combining World Model Forecasting and Vision-Language Reasoning
- 用视觉语言模型生成行为指令,引导世界模型进行驾驶决策。
- 在真实场景数据集上实现更准确的环境动态预测,提升驾驶策略可解释性。
- 适合研究自动驾驶感知与决策融合、多模态模型设计的学者参考。
自动驾驶系统依赖于能够理解高层场景上下文并精确预测周围环境动态的模型。视觉-语言模型(VLM)近期成为决策与场景理解的有力工具,具备强大的上下文推理能力。然而,其有限的空间理解能力限制了其作为端到端驾驶模型的有效性。世界模型(WM)通过内化环境动态来预测未来场景演变,被探索用于自车运动预测和自动驾驶的基础模型,是解决该领域关键挑战的有前景方向,尤其在增强泛化能力的同时保持动态预测性能。为融合基于上下文的决策与预测优势,我们提出WorldVLM:一种将VLM与WM统一的混合架构。在设计中,高阶VLM生成行为指令以引导驾驶用的WM,实现可解释且情境感知的动作。我们评估了多种条件策略,并提供了对混合设计挑战的见解。
原文摘要 · Abstract (English)
Autonomous driving systems depend on on models that can reason about high-level scene contexts and accurately predict the dynamics of their surrounding environment. Vision- Language Models (VLMs) have recently emerged as promising tools for decision-making and scene understanding, offering strong capabilities in contextual reasoning. However, their limited spatial comprehension constrains their effectiveness as end-to-end driving models. World Models (WM) internalize environmental dynamics to predict future scene evolution. Recently explored as ego-motion predictors and foundation models for autonomous driving, they represent a promising direction for addressing key challenges in the field, particularly enhancing generalization while maintaining dynamic prediction. To leverage the complementary strengths of context-based decision making and prediction, we propose WorldVLM: A hybrid architecture that unifies VLMs and WMs. In our design, the high-level VLM generates behavior commands to guide the driving WM, enabling interpretable and context-aware actions. We evaluate conditioning strategies and provide insights into the hybrid design challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。