arXiv:2507.10672cs.ROcs.CV2025-07综述被引 49

系统梳理视觉语言动作模型在机器人操作中的研究进展。

Vision Language Action Models in Robotic Manipulation: A Systematic Review

  • 分析102个VLA模型、26个数据集与12个仿真平台的架构与应用。
  • 提出二维数据表征框架,揭示当前数据在语义与模态对齐上的空白区。
  • 为通用机器人智能提供从训练到部署的路线图,适合研究者与工业开发者。

视觉语言动作(VLA)模型代表了机器人领域的范式转变,旨在将视觉感知、自然语言理解与具身控制统一在一个学习框架中。本文系统综述了102个VLA模型、26个基础数据集和12个仿真平台,全面剖析其在机器人操作与指令驱动自主中的发展与评估。通过任务复杂度、模态多样性与数据规模的新评价标准,对比分析各数据集对通用策略学习的适用性。提出二维数据表征框架,基于语义丰富度与多模态对齐程度,揭示当前数据布局中的未探索区域。仿真环境评估聚焦大规模数据生成能力、仿真到现实的迁移效率及任务多样性。结合学术与产业贡献,识别现存挑战并提出可扩展预训练、模块化架构设计与鲁棒多模态对齐等战略方向。本综述为具身智能与机器人控制提供技术参考与概念路线图,涵盖数据生成至真实世界部署的全链条洞察。

原文摘要 · Abstract (English)

Vision Language Action (VLA) models represent a transformative shift in robotics, with the aim of unifying visual perception, natural language understanding, and embodied control within a single learning framework. This review presents a comprehensive and forward-looking synthesis of the VLA paradigm, with a particular emphasis on robotic manipulation and instruction-driven autonomy. We comprehensively analyze 102 VLA models, 26 foundational datasets, and 12 simulation platforms that collectively shape the development and evaluation of VLAs models. These models are categorized into key architectural paradigms, each reflecting distinct strategies for integrating vision, language, and control in robotic systems. Foundational datasets are evaluated using a novel criterion based on task complexity, variety of modalities, and dataset scale, allowing a comparative analysis of their suitability for generalist policy learning. We introduce a two-dimensional characterization framework that organizes these datasets based on semantic richness and multimodal alignment, showing underexplored regions in the current data landscape. Simulation environments are evaluated for their effectiveness in generating large-scale data, as well as their ability to facilitate transfer from simulation to real-world settings and the variety of supported tasks. Using both academic and industrial contributions, we recognize ongoing challenges and outline strategic directions such as scalable pretraining protocols, modular architectural design, and robust multimodal alignment strategies. This review serves as both a technical reference and a conceptual roadmap for advancing embodiment and robotic control, providing insights that span from dataset generation to real world deployment of generalist robotic agents.

机器人多模态大模型泛化智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。