arXiv:2509.19012cs.ROcs.AI2025-09综述被引 56

综述视觉语言动作模型,梳理机器人智能新范式

Pure Vision Language Action (VLA) Models: A Comprehensive Survey

  • 按生成方式分类:自回归、扩散、强化学习等五类方法
  • 涵盖300+研究,系统分析现有VLA模型与应用场景
  • 适合关注通用机器人、多模态智能的科研人员

视觉语言动作(VLA)模型的兴起标志着从传统策略控制向通用机器人范式的转变,将视觉语言模型从被动序列生成器重构为在复杂动态环境中进行操作与决策的主动代理。本综述深入探讨先进的VLA方法,旨在提供清晰的分类体系和系统的全面回顾。文章对VLA应用在不同场景中的表现进行了综合分析,并将其方法分为自回归、扩散、强化学习、混合及专用五类,详述其动机、核心策略与实现方式。同时介绍了基础数据集、评估基准与仿真平台。基于当前研究现状,进一步提出关键挑战与未来方向,推动VLA模型与可泛化机器人研究的发展。通过整合超过三百篇近期研究,该综述勾勒出这一快速演进领域的轮廓,凸显塑造可扩展、通用VLA方法的机会与挑战。

原文摘要 · Abstract (English)

The emergence of Vision Language Action (VLA) models marks a paradigm shift from traditional policy-based control to generalized robotics, reframing Vision Language Models (VLMs) from passive sequence generators into active agents for manipulation and decision-making in complex, dynamic environments. This survey delves into advanced VLA methods, aiming to provide a clear taxonomy and a systematic, comprehensive review of existing research. It presents a comprehensive analysis of VLA applications across different scenarios and classifies VLA approaches into several paradigms: autoregression-based, diffusion-based, reinforcement-based, hybrid, and specialized methods; while examining their motivations, core strategies, and implementations in detail. In addition, foundational datasets, benchmarks, and simulation platforms are introduced. Building on the current VLA landscape, the review further proposes perspectives on key challenges and future directions to advance research in VLA models and generalizable robotics. By synthesizing insights from over three hundred recent studies, this survey maps the contours of this rapidly evolving field and highlights the opportunities and challenges that will shape the development of scalable, general-purpose VLA methods.

VLA模型机器人智能多模态综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。