arXiv:2512.04441cs.CV2025-12被引 11

融合世界模型与视觉语言模型,实现端到端自动驾驶的智能决策与轨迹生成。

MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving

  • 通过未来模拟生成前瞻性轨迹候选,结合环境感知进行条件化推理。
  • 在NAVSIM-v1/v2上实现多维指标最优,安全性和通用性显著提升。
  • 适合关注可解释性、类人驾驶的自动驾驶研究者与开发者。

端到端自动驾驶(E2E-AD)已成为新范式,其中轨迹规划至关重要。现有方法主要分两类:以轨迹生成为导向,注重高质量轨迹输出但决策机制简单;以轨迹选择为导向,虽能多维度评估但生成能力不足。本文提出MindDrive框架,融合高质量轨迹生成与全面决策推理,构建“上下文模拟—候选生成—多目标权衡”的结构化推理范式。其中,基于世界动作模型(WaM)的前瞻轨迹生成器(FaTG)实现自车条件下的‘假如’情景模拟,预测潜在未来场景并生成前瞻性轨迹。在此基础上,视觉语言模型导向的评估器(VLoE)利用大模型的推理能力,在安全性、舒适性和效率三方面进行多目标评估,实现符合人类认知的决策。在NAVSIM-v1和NAVSIM-v2基准上的大量实验表明,MindDrive在多维驾驶指标上达到当前最优表现,显著提升安全性、合规性与泛化能力。该工作为可解释、认知引导的自动驾驶提供了可行路径。

原文摘要 · Abstract (English)

End-to-End autonomous driving (E2E-AD) has emerged as a new paradigm, where trajectory planning plays a crucial role. Existing studies mainly follow two directions: trajectory generation oriented, which focuses on producing high-quality trajectories with simple decision mechanisms, and trajectory selection oriented, which performs multi-dimensional evaluation to select the best trajectory yet lacks sufficient generative capability. In this work, we propose MindDrive, a harmonized framework that integrates high-quality trajectory generation with comprehensive decision reasoning. It establishes a structured reasoning paradigm of "context simulation - candidate generation - multi-objective trade-off". In particular, the proposed Future-aware Trajectory Generator (FaTG), based on a World Action Model (WaM), performs ego-conditioned "what-if" simulations to predict potential future scenes and generate foresighted trajectory candidates. Building upon this, the VLM-oriented Evaluator (VLoE) leverages the reasoning capability of a large vision-language model to conduct multi-objective evaluations across safety, comfort, and efficiency dimensions, leading to reasoned and human-aligned decision making. Extensive experiments on the NAVSIM-v1 and NAVSIM-v2 benchmarks demonstrate that MindDrive achieves state-of-the-art performance across multi-dimensional driving metrics, significantly enhancing safety, compliance, and generalization. This work provides a promising path toward interpretable and cognitively guided autonomous driving.

自动驾驶世界模型视觉语言模型轨迹生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。