提出高效通用机器人视觉语言动作模型,显著提升真实场景操作能力
What Matters in Building Vision-Language-Action Models for Generalist Robots
- 基于8种大模型、4类策略架构,系统验证关键设计选择
- 在3个仿真与真实任务中达到新基准性能,减少人工设计依赖
- 开源完整框架与训练方案,支持快速集成与实验复现
为将基础视觉语言模型(VLM)应用于机器人任务与运动规划,学界提出了多种向VLM注入动作组件的方法,构建视觉-语言-动作模型(VLA)。本文揭示了影响VLA在机器人操作任务中性能的关键因素,聚焦三个核心设计选择:选用何种骨干网络、如何构建VLA架构、何时引入跨具身数据。通过超过8种VLM骨干、4种策略架构、600余组不同实验的系统性研究,我们验证了优选方案并开发出新一代无须大量人工设计的RoboVLMs模型,在三个仿真任务和真实世界实验中均取得当前最优表现。同时,我们公开了高度灵活的RoboVLMs框架,支持新VLM轻松接入与自由组合设计选项,并提供完整代码、模型、数据集及训练评估手册,详见 robovlms.github.io。
原文摘要 · Abstract (English)
To utilize Foundation Vision Language Models (VLMs) for robotic tasks and motion planning, the community has proposed different methods for injecting action components into VLMs and building the Vision-Language-Action models (VLAs). In this work, we disclose the key factors that significantly influence the performance of VLA on robot manipulation problems and focus on answering three essential design choices: which backbone to select, how to formulate the VLA architectures, and when to add cross-embodiment data. The obtained results convince us firmly to explain why we prefer VLA and develop a new family of VLAs, RoboVLMs, which require very few manual designs and achieve a new state-of-the-art performance in three simulation tasks and real-world experiments. Through our extensive experiments, which include over 8 VLM backbones, 4 policy architectures, and over 600 distinct designed experiments, we provide a detailed guidebook for the future design of VLAs. In addition to the study, the highly flexible RoboVLMs framework, which supports easy integrations of new VLMs and free combinations of various design choices, is made public to facilitate future research. We open-source all details, including codes, models, datasets, and toolkits, along with detailed training and evaluation recipes at: robovlms.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。