arXiv:2607.06706cs.ROcs.AI2026-07综述被引 1

VLA模型让机器人听懂指令并执行复杂动作,跨越双臂操作与无人机控制。

Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review

论文配图:Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review
图 1 · 摘自论文原文
  • 统一视觉、语言与动作生成的单模型框架,直接从图像理解指令
  • 183篇论文梳理出跨领域通用的协调策略与训练方法
  • 适合研究机器人智能、多模态系统或具身智能的学者参考

视觉语言动作(VLA)模型将视觉感知、自然语言理解和动作生成整合到单一基础模型中,使机器人能直接从摄像头图像理解并执行如‘折叠毛巾’或‘飞向红房子’等指令。由于继承了互联网规模预训练的世界知识,VLA已成为基于学习的机器人操作主流框架,其中双臂协同是最高挑战:每只手臂具7个自由度,需协同完成折叠、组装和重定向物体。无人机自主飞行面临类似难题:在严格延迟和载荷限制下,必须根据视觉观测协调推力、姿态甚至机械臂指令。本综述涵盖2017至2026年间183项工作,按七维结构组织:VLA架构、训练方法、动作表示、双臂协作(2022–2026)、无人机导航与控制(2017–2026)、语言定位以及记忆与世界模型等共性问题。结果表明,双臂VLA所发展的协调策略、训练方法与动作表示可有效迁移至无人机系统,并识别出两个领域共有的14个研究方向。

原文摘要 · Abstract (English)

Vision Language Action (VLA) models unify visual perception, natural-language understanding, and action generation within a single foundation model, allowing a robot to follow instructions such as fold the towel or fly to the red building directly from camera images. Because VLAs inherit world knowledge from internet-scale pre-training, they have become the dominant framework for learning-based manipulation, with bimanual coordination serving as the most demanding testbed: two arms with 7 degrees of freedom each must move in concert to fold, assemble, and reorient objects. Unmanned aerial robotics faces a structurally similar challenge: a drone must coordinate thrust, attitude, and increasingly gripper commands from visual observations under strict latency and payload constraints. This review covers 183 contributions spanning 2017-2026 and organized along seven dimensions: VLA architectures, training recipes, action representations, bimanual coordination (2022-2026), unmanned aerial vehicle (UAV) navigation and control (2017-2026), language grounding, and cross-cutting concerns including memory and world models. We show that the coordination strategies, training recipes, and action representations developed for bimanual VLAs transfer to unmanned aerial systems and identify fourteen research directions across both domains.

VLA模型机器人控制双臂协作无人机导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。