综述视觉-语言-动作模型在机器人操作中的进展与挑战
Survey of Vision-Language-Action Models for Embodied Manipulation
- 梳理VLA模型架构演进路径,涵盖五大核心维度
- 分析训练数据、预训练与后训练方法对性能的影响
- 适合关注具身智能机器人控制的研究者和工程师
具身智能系统通过持续与环境交互提升智能体能力,受到学界与产业界广泛关注。视觉-语言-动作(VLA)模型受大基础模型发展启发,作为通用机器人控制框架,显著增强了具身智能系统中智能体与环境的交互能力,拓展了具身AI机器人的应用范围。本综述全面回顾了面向具身操作的VLA模型研究。首先,梳理了VLA架构的发展历程;其次,从五个关键维度展开深入分析:模型结构、训练数据集、预训练方法、后训练方法及模型评估;最后,总结了VLA发展与实际部署中的核心挑战,并展望了未来有前景的研究方向。
原文摘要 · Abstract (English)
Embodied intelligence systems, which enhance agent capabilities through continuous environment interactions, have garnered significant attention from both academia and industry. Vision-Language-Action models, inspired by advancements in large foundation models, serve as universal robotic control frameworks that substantially improve agent-environment interaction capabilities in embodied intelligence systems. This expansion has broadened application scenarios for embodied AI robots. This survey comprehensively reviews VLA models for embodied manipulation. Firstly, it chronicles the developmental trajectory of VLA architectures. Subsequently, we conduct a detailed analysis of current research across 5 critical dimensions: VLA model structures, training datasets, pre-training methods, post-training methods, and model evaluation. Finally, we synthesize key challenges in VLA development and real-world deployment, while outlining promising future research directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。