用视觉语言动作流模型实现通用机器人控制,让机器人听懂指令并自主完成复杂任务。
$π_0$: A Vision-Language-Action Flow Model for General Robot Control
- 基于预训练视觉语言模型构建流匹配架构,继承互联网级语义知识。
- 在多平台机器人数据上训练,零样本直接执行洗衣、清洁、组装等任务。
- 支持指令理解与微调学习新技能,适合研究通用机器人系统者参考。
机器人学习有望释放灵活、通用且灵巧机器人系统的全部潜力,并解决人工智能中一些深层次问题。然而,将机器人学习推向真实世界系统所需的通用性水平,面临数据、泛化性和鲁棒性方面的重大挑战。本文探讨了通用机器人策略(即机器人基础模型)如何应对这些挑战,并提出一种基于预训练视觉语言模型(VLM)的新型流匹配架构,以继承互联网规模的语义知识。该模型在来自多种灵巧机器人平台的大规模多样化数据集上进行训练,涵盖单臂机器人、双臂机器人和移动操作机器人。评估表明,模型在预训练后具备零样本执行能力,可理解人类语言指令及高层VLM策略指令,并通过微调获得新技能。实验覆盖洗衣折叠、桌面清洁、盒子组装等多种任务。
原文摘要 · Abstract (English)
Robot learning holds tremendous promise to unlock the full potential of flexible, general, and dexterous robot systems, as well as to address some of the deepest questions in artificial intelligence. However, bringing robot learning to the level of generality required for effective real-world systems faces major obstacles in terms of data, generalization, and robustness. In this paper, we discuss how generalist robot policies (i.e., robot foundation models) can address these challenges, and how we can design effective generalist robot policies for complex and highly dexterous tasks. We propose a novel flow matching architecture built on top of a pre-trained vision-language model (VLM) to inherit Internet-scale semantic knowledge. We then discuss how this model can be trained on a large and diverse dataset from multiple dexterous robot platforms, including single-arm robots, dual-arm robots, and mobile manipulators. We evaluate our model in terms of its ability to perform tasks in zero shot after pre-training, follow language instructions from people and from a high-level VLM policy, and its ability to acquire new skills via fine-tuning. Our results cover a wide variety of tasks, such as laundry folding, table cleaning, and assembling boxes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。