系统梳理视觉-语言-动作模型的模块、里程碑与核心挑战
An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- 按感知-执行-泛化路径拆解VLA模型核心模块
- 归纳五大关键挑战:表征、执行、泛化、安全与评估
- 适合机器人、AI具身智能领域研究者快速入门
视觉-语言-动作(VLA)模型正推动机器人领域的变革,使机器能够理解指令并交互物理世界。该领域模型与数据集迅猛发展,既令人振奋又难跟进。本文作为综述,为研究者提供清晰的结构化指南:从基础模块出发,梳理关键里程碑,深入探讨当前前沿的核心挑战。主要贡献在于系统分解五大挑战:(1)表征,(2)执行,(3)泛化,(4)安全,(5)数据集与评估。该结构对应通用智能体的发展路线:建立基础感知-动作回路,跨模态与环境扩展能力,最终实现可信部署,均依赖于数据基础设施支撑。每项挑战下回顾现有方法并指明未来方向。本文既是新入者的奠基读物,也是资深研究者的战略蓝图,旨在加速学习并激发新思路。项目页面持续更新:https://suyuz1.github.io/VLA-Survey-Anatomy/
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models are driving a revolution in robotics, enabling machines to understand instructions and interact with the physical world. This field is exploding with new models and datasets, making it both exciting and challenging to keep pace with. This survey offers a clear and structured guide to the VLA landscape. We design it to follow the natural learning path of a researcher: we start with the basic Modules of any VLA model, trace the history through key Milestones, and then dive deep into the core Challenges that define recent research frontier. Our main contribution is a detailed breakdown of the five biggest challenges in: (1) Representation, (2) Execution, (3) Generalization, (4) Safety, and (5) Dataset and Evaluation. This structure mirrors the developmental roadmap of a generalist agent: establishing the fundamental perception-action loop, scaling capabilities across diverse embodiments and environments, and finally ensuring trustworthy deployment-all supported by the essential data infrastructure. For each of them, we review existing approaches and highlight future opportunities. We position this paper as both a foundational guide for newcomers and a strategic roadmap for experienced researchers, with the dual aim of accelerating learning and inspiring new ideas in embodied intelligence. A live version of this survey, with continuous updates, is maintained on our \href{https://suyuz1.github.io/VLA-Survey-Anatomy/}{project page}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。