arXiv:2604.02190cs.CVcs.RO2026-04被引 14

统一视觉语言动作模型,解决自动驾驶感知与推理的矛盾

UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving

  • 用混合专家结构分离感知、理解与规划模块
  • 在nuScenes和Bench2Drive上达领先性能
  • 适合需要多任务协同的自动驾驶系统研究者

视觉-语言-动作(VLA)模型在自动驾驶中崭露头角,有望利用丰富的世界知识提升系统认知能力。然而,现有系统在空间感知与语义推理之间面临关键困境:直接使用2D视觉语言模型导致空间感知不足,而引入3D表示又削弱了原有推理能力。本文认为该困境源于感知与推理在共享参数中的耦合优化。为此,提出UniDriveVLA,一种基于混合变压器的统一驾驶视觉-语言-动作模型,通过专家解耦解决感知-推理冲突。其包含三个专家:驾驶理解、场景感知和行为规划,由掩码联合注意力协调。结合稀疏感知范式与三阶段渐进训练策略,在保持语义推理能力的同时提升空间感知。大量实验表明,UniDriveVLA在nuScenes开环评估和Bench2Drive闭环评估中均达到顶尖水平,并在3D检测、在线建图、运动预测和驾驶导向VQA等任务上表现优异,展现作为统一自动驾驶模型的广泛适用性。代码与模型已开源。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have recently emerged in autonomous driving, with the promise of leveraging rich world knowledge to improve the cognitive capabilities of driving systems. However, adapting such models for driving tasks currently faces a critical dilemma between spatial perception and semantic reasoning. Consequently, existing VLA systems are forced into suboptimal compromises: directly adopting 2D Vision-Language Models yields limited spatial perception, whereas enhancing them with 3D spatial representations often impairs the native reasoning capacity of VLMs. We argue that this dilemma largely stems from the coupled optimization of spatial perception and semantic reasoning within shared model parameters. To overcome this, we propose UniDriveVLA, a Unified Driving Vision-Language-Action model based on Mixture-of-Transformers that addresses the perception-reasoning conflict via expert decoupling. Specifically, it comprises three experts for driving understanding, scene perception, and action planning, which are coordinated through masked joint attention. In addition, we combine a sparse perception paradigm with a three-stage progressive training strategy to improve spatial perception while maintaining semantic reasoning capability. Extensive experiments show that UniDriveVLA achieves state-of-the-art performance in open-loop evaluation on nuScenes and closed-loop evaluation on Bench2Drive. Moreover, it demonstrates strong performance across a broad range of perception, prediction, and understanding tasks, including 3D detection, online mapping, motion forecasting, and driving-oriented VQA, highlighting its broad applicability as a unified model for autonomous driving. Code and model have been released at https://github.com/xiaomi-research/unidrivevla

自动驾驶视觉语言多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。