arXiv:2503.23463cs.CV2025-03AAAI被引 206

用大模型实现端到端自动驾驶,能理解指令并规划路径。

OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model

  • 融合视觉、车辆状态和语言指令,生成带空间定位的驾驶动作。
  • 在nuScenes数据集上轨迹预测与问答任务均达顶尖水平。
  • 擅长理解复杂指令,在挑战场景下仍能生成合理路径。

我们提出OpenDriveVLA,一种基于开源大语言模型的视觉-语言-动作模型,用于端到端自动驾驶。该模型通过融合2D与3D实例感知视觉表示、自车状态及语言指令,生成具有空间定位的驾驶动作。为弥合视觉表征与语言嵌入之间的模态差异,我们引入分层视觉-语言对齐机制,将2D与3D结构化视觉标记投影至统一语义空间。此外,将结构化智能体-环境-自车交互建模融入自回归解码过程,使模型能够捕捉细粒度空间依赖关系与行为感知动态,这对可靠轨迹规划至关重要。在nuScenes数据集上的大量实验表明,OpenDriveVLA在开环轨迹规划与驾驶相关问答任务中均达到当前最优性能。定性分析进一步展示其在复杂场景下遵循高层驾驶指令并生成合理轨迹的能力,凸显其在下一代端到端自动驾驶中的潜力。

原文摘要 · Abstract (English)

We present OpenDriveVLA, a Vision Language Action model designed for end-to-end autonomous driving, built upon open-source large language models. OpenDriveVLA generates spatially grounded driving actions by leveraging multimodal inputs, including 2D and 3D instance-aware visual representations, ego vehicle states, and language commands. To bridge the modality gap between driving visual representations and language embeddings, we introduce a hierarchical vision language alignment process, projecting both 2D and 3D structured visual tokens into a unified semantic space. Furthermore, we incorporate structured agent environment ego interaction modeling into the autoregressive decoding process, enabling the model to capture fine-grained spatial dependencies and behavior-aware dynamics critical for reliable trajectory planning. Extensive experiments on the nuScenes dataset demonstrate that OpenDriveVLA achieves state-of-the-art results across open-loop trajectory planning and driving-related question answering tasks. Qualitative analyses further illustrate its capability to follow high-level driving commands and generate trajectories under challenging scenarios, highlighting its potential for next-generation end-to-end autonomous driving.

端到端驾驶多模态大模型轨迹规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。