arXiv:2510.07077cs.ROcs.AI2025-10中稿 · IEEE Access, websi…综述被引 151

整合视觉、语言与动作的模型让机器人更灵活地应对真实世界任务。

Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications

论文配图:Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
图 1 · 摘自论文原文
  • 统一处理视觉、语言和动作数据,实现跨任务泛化。
  • 覆盖从架构到硬件平台的完整系统设计,支持实际部署。
  • 适合想落地机器人应用的研究者和工程师参考。

随着大型语言模型(LLMs)和视觉-语言模型(VLMs)在机器人领域的应用日益广泛,视觉-语言-动作(VLA)模型近年来受到广泛关注。通过大规模统一视觉、语言和动作数据——这些以往被分别研究的模态——VLA模型旨在学习可泛化于不同任务、物体、机器人本体和环境的策略。这种泛化能力有望使机器人在无需或仅需极少额外任务特定数据的情况下,解决新任务,从而推动更灵活、可扩展的真实世界部署。与以往聚焦于动作表示或高层架构的综述不同,本文提供了一个全面的全栈式综述,涵盖软件与硬件组件。具体包括:VLA的策略与架构演进、核心架构与模块、模态特异性处理技术以及学习范式。此外,为支持真实机器人应用中的部署,本文还回顾了常用机器人平台、数据采集策略、公开数据集、数据增强方法及评估基准。整篇综述旨在为机器人社区提供将VLA应用于实际系统的实用指导。所有参考文献按训练方式、评估方法、模态和数据集分类,详见项目网站:https://vla-survey.github.io。

原文摘要 · Abstract (English)

Amid growing efforts to leverage advances in large language models (LLMs) and vision-language models (VLMs) for robotics, Vision-Language-Action (VLA) models have recently gained significant attention. By unifying vision, language, and action data at scale, which have traditionally been studied separately, VLA models aim to learn policies that generalise across diverse tasks, objects, embodiments, and environments. This generalisation capability is expected to enable robots to solve novel downstream tasks with minimal or no additional task-specific data, facilitating more flexible and scalable real-world deployment. Unlike previous surveys that focus narrowly on action representations or high-level model architectures, this work offers a comprehensive, full-stack review, integrating both software and hardware components of VLA systems. In particular, this paper provides a systematic review of VLAs, covering their strategy and architectural transition, architectures and building blocks, modality-specific processing techniques, and learning paradigms. In addition, to support the deployment of VLAs in real-world robotic applications, we also review commonly used robot platforms, data collection strategies, publicly available datasets, data augmentation methods, and evaluation benchmarks. Throughout this comprehensive survey, this paper aims to offer practical guidance for the robotics community in applying VLAs to real-world robotic systems. All references categorized by training approach, evaluation method, modality, and dataset are available in the table on our project website: https://vla-survey.github.io .

机器人多模态大模型泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。