arXiv:2505.04769cs.CV2025-05被引 95

整合视觉、语言与行动的智能模型,让机器人更懂世界也更会做事。

Vision-Language-Action (VLA) Models: Concepts, Progress, Applications and Challenges

  • 统一视觉、语言与动作的跨模态框架,支持复杂任务理解与执行
  • 涵盖80多个最新模型,覆盖自动驾驶到人形机器人等多领域应用
  • 提出通用智能体发展路径,适合研究机器人与通用AI的开发者

视觉-语言-动作(VLA)模型代表人工智能的重大进展,旨在将感知、自然语言理解和具身行动统一于单一计算框架中。本文系统综述过去三年超过80个VLA模型的进展,围绕五大主题展开:概念基础、架构创新、高效训练策略、实时推理加速及多样化应用场景。涵盖自动驾驶、医疗与工业机器人、精准农业、人形机器人和增强现实等领域。分析当前挑战并提出代理适应与跨具身规划等解决方案。展望未来,VLA模型、视觉语言模型(VLMs)与代理型AI将融合,推动社会对齐、可适应、通用的具身智能体发展。项目代码库已开源:https://github.com/Applied-AI-Research-Lab/Vision-Language-Action-Models-Concepts-Progress-Applications-and-Challenges。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models mark a transformative advancement in artificial intelligence, aiming to unify perception, natural language understanding, and embodied action within a single computational framework. This foundational review presents a comprehensive synthesis of recent advancements in Vision-Language-Action models, systematically organized across five thematic pillars that structure the landscape of this rapidly evolving field. We begin by establishing the conceptual foundations of VLA systems, tracing their evolution from cross-modal learning architectures to generalist agents that tightly integrate vision-language models (VLMs), action planners, and hierarchical controllers. Our methodology adopts a rigorous literature review framework, covering over 80 VLA models published in the past three years. Key progress areas include architectural innovations, efficient training strategies, and real-time inference accelerations. We explore diverse application domains such as autonomous vehicles, medical and industrial robotics, precision agriculture, humanoid robotics, and augmented reality. We analyzed challenges and propose solutions including agentic adaptation and cross-embodiment planning. Furthermore, we outline a forward-looking roadmap where VLA models, VLMs, and agentic AI converge to strengthen socially aligned, adaptive, and general-purpose embodied agents. This work, is expected to serve as a foundational reference for advancing intelligent, real-world robotics and artificial general intelligence. The project repository is available on GitHub as https://github.com/Applied-AI-Research-Lab/Vision-Language-Action-Models-Concepts-Progress-Applications-and-Challenges. [Index Terms: Vision Language Action, VLA, Vision Language Models, VLMs, Action Tokenization, NLP]

VLA模型机器人通用智能跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。