首份自动驾驶视觉语言动作模型综述,梳理技术演进与评测标准
A Survey on Vision-Language-Action Models for Autonomous Driving
- 系统拆解VLA模型的核心架构组件
- 对比20余款模型在安全、准确与可解释性上的表现
- 适合关注自动驾驶智能决策与可解释性的研究者
多模态大模型的快速发展推动了视觉-语言-动作(VLA)范式的兴起,该范式将视觉感知、自然语言理解与控制策略整合于统一策略中。自动驾驶领域研究人员正积极适配此类方法。这类模型有望使自动驾驶车辆能够理解高层指令、推理复杂交通场景并自主决策。然而现有文献分散且快速扩张。本文首次全面综述自动驾驶视觉语言动作模型(VLA4AD),(i) 形式化归纳近期工作的共性架构组件,(ii) 追溯从早期解释型到以推理为核心的VLA模型演变,(iii) 对超过20个代表性模型进行对比分析。同时整合现有数据集与基准,突出联合评估驾驶安全、准确性与可解释性的评测协议。最后,详述鲁棒性、实时效率与形式化验证等开放挑战,并展望未来方向。本综述为推进可解释且社会对齐的自动驾驶系统提供简洁完整的参考。GitHub仓库见:https://github.com/JohnsonJiang1996/Awesome-VLA4AD
原文摘要 · Abstract (English)
The rapid progress of multimodal large language models (MLLM) has paved the way for Vision-Language-Action (VLA) paradigms, which integrate visual perception, natural language understanding, and control within a single policy. Researchers in autonomous driving are actively adapting these methods to the vehicle domain. Such models promise autonomous vehicles that can interpret high-level instructions, reason about complex traffic scenes, and make their own decisions. However, the literature remains fragmented and is rapidly expanding. This survey offers the first comprehensive overview of VLA for Autonomous Driving (VLA4AD). We (i) formalize the architectural building blocks shared across recent work, (ii) trace the evolution from early explainer to reasoning-centric VLA models, and (iii) compare over 20 representative models according to VLA's progress in the autonomous driving domain. We also consolidate existing datasets and benchmarks, highlighting protocols that jointly measure driving safety, accuracy, and explanation quality. Finally, we detail open challenges - robustness, real-time efficiency, and formal verification - and outline future directions of VLA4AD. This survey provides a concise yet complete reference for advancing interpretable socially aligned autonomous vehicles. Github repo is available at \href{https://github.com/JohnsonJiang1996/Awesome-VLA4AD}{SicongJiang/Awesome-VLA4AD}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。