综述大模型驱动的机器人操作视觉语言动作系统,梳理架构与前沿方向。
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- 按一体化与分层架构分类,系统梳理VLA模型设计范式
- 整合强化学习、人类视频学习等多领域技术提升泛化能力
- 适合关注具身智能、机器人通用控制的研究者参考
机器人操作是机器人学与具身人工智能的关键前沿,需精确运动控制与多模态理解,但传统规则方法难以在非结构化新环境中扩展或泛化。近年来,基于大规模视觉-语言模型(VLM)预训练的视觉-语言-动作(VLA)模型成为变革性范式。本综述首次系统性、分类导向地回顾了基于大VLM的VLA模型在机器人操作中的应用。我们首先明确定义大VLM-based VLA模型,并划分两大主要架构范式:(1) 单体模型,包含单系统与双系统设计,集成程度不同;(2) 分层模型,通过可解释中间表示显式解耦规划与执行。在此基础上,深入分析大VLM-based VLA模型:(1) 与强化学习、免训练优化、从人类视频学习、世界模型融合等先进领域的结合;(2) 整合其架构特征、运行优势及支撑其发展的数据集与基准;(3) 指出未来方向,包括记忆机制、4D感知、高效适应、多智能体协作等新兴能力。本综述整合近期进展,弥合现有分类不一致问题,缓解研究碎片化,填补大VLM与机器人操作交叉研究的关键空白。我们提供持续更新的项目页面以追踪进展:https://github.com/JiuTian-VL/Large-VLM-based-VLA-for-Robotic-Manipulation
原文摘要 · Abstract (English)
Robotic manipulation, a key frontier in robotics and embodied AI, requires precise motor control and multimodal understanding, yet traditional rule-based methods fail to scale or generalize in unstructured, novel environments. In recent years, Vision-Language-Action (VLA) models, built upon Large Vision-Language Models (VLMs) pretrained on vast image-text datasets, have emerged as a transformative paradigm. This survey provides the first systematic, taxonomy-oriented review of large VLM-based VLA models for robotic manipulation. We begin by clearly defining large VLM-based VLA models and delineating two principal architectural paradigms: (1) monolithic models, encompassing single-system and dual-system designs with differing levels of integration; and (2) hierarchical models, which explicitly decouple planning from execution via interpretable intermediate representations. Building on this foundation, we present an in-depth examination of large VLM-based VLA models: (1) integration with advanced domains, including reinforcement learning, training-free optimization, learning from human videos, and world model integration; (2) synthesis of distinctive characteristics, consolidating architectural traits, operational strengths, and the datasets and benchmarks that support their development; (3) identification of promising directions, including memory mechanisms, 4D perception, efficient adaptation, multi-agent cooperation, and other emerging capabilities. This survey consolidates recent advances to resolve inconsistencies in existing taxonomies, mitigate research fragmentation, and fill a critical gap through the systematic integration of studies at the intersection of large VLMs and robotic manipulation. We provide a regularly updated project page to document ongoing progress: https://github.com/JiuTian-VL/Large-VLM-based-VLA-for-Robotic-Manipulation
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。