arXiv:2512.16760cs.RO2025-12被引 38

让自动驾驶从分步执行转向语言驱动的智能决策。

Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future

  • 用视觉-语言-动作统一模型替代传统分步流程
  • 支持指令理解与复杂场景下的推理决策
  • 适合研究通用自动驾驶系统与人机交互的学者

自动驾驶长期依赖“感知-决策-执行”的模块化流水线,其手工设计的接口和规则组件在复杂或长尾场景中常失效,级联结构还导致感知误差传播,影响下游规划与控制。视觉-动作(VA)模型虽通过直接学习视觉到动作的映射缓解部分问题,但仍存在黑箱性、对分布偏移敏感、缺乏结构化推理和指令遵循能力等缺陷。近期大语言模型(LLMs)与多模态学习的发展推动了视觉-语言-动作(VLA)框架的兴起,该框架将感知与语言引导的决策融合。通过整合视觉理解、语言推理与可执行输出,VLAs为更可解释、泛化更强且符合人类意图的驾驶策略提供了路径。本文系统梳理了自动驾驶领域VLA的发展脉络,从早期VA方法演进至现代VLA框架,将现有方法归纳为两大范式:端到端VLA(将感知、推理与规划统一于单个模型)与双系统VLA(将慢速推理由视觉语言模型完成,快速安全执行由规划器实现)。进一步区分文本/数值动作生成器以及显式/隐式引导机制。总结代表性数据集与评估基准,并指出鲁棒性、可解释性与指令保真度等关键挑战与未来方向。整体目标是为构建以人为本的自动驾驶系统奠定坚实基础。

原文摘要 · Abstract (English)

Autonomous driving has long relied on modular "Perception-Decision-Action" pipelines, where hand-crafted interfaces and rule-based components often break down in complex or long-tailed scenarios. Their cascaded design further propagates perception errors, degrading downstream planning and control. Vision-Action (VA) models address some limitations by learning direct mappings from visual inputs to actions, but they remain opaque, sensitive to distribution shifts, and lack structured reasoning or instruction-following capabilities. Recent progress in Large Language Models (LLMs) and multimodal learning has motivated the emergence of Vision-Language-Action (VLA) frameworks, which integrate perception with language-grounded decision making. By unifying visual understanding, linguistic reasoning, and actionable outputs, VLAs offer a pathway toward more interpretable, generalizable, and human-aligned driving policies. This work provides a structured characterization of the emerging VLA landscape for autonomous driving. We trace the evolution from early VA approaches to modern VLA frameworks and organize existing methods into two principal paradigms: End-to-End VLA, which integrates perception, reasoning, and planning within a single model, and Dual-System VLA, which separates slow deliberation (via VLMs) from fast, safety-critical execution (via planners). Within these paradigms, we further distinguish subclasses such as textual vs. numerical action generators and explicit vs. implicit guidance mechanisms. We also summarize representative datasets and benchmarks for evaluating VLA-based driving systems and highlight key challenges and open directions, including robustness, interpretability, and instruction fidelity. Overall, this work aims to establish a coherent foundation for advancing human-compatible autonomous driving systems.

自动驾驶多模态大模型智能决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。