GR-3是能泛化处理新物体与复杂指令的通用机器人模型。
GR-3 Technical Report

- 基于视觉-语言-动作协同训练,融合网络数据与人类轨迹微调
- 仅需少量人类示范即可高效适配新任务,实测优于现有基线方法π₀
- 适配双臂移动机器人ByteMini,可完成长周期精细操作任务
本文报告了构建通用机器人策略的最新进展——GR-3。GR-3是一个大规模视觉-语言-动作(VLA)模型,具备在新物体、新环境及包含抽象概念的指令下良好泛化的能力。该模型可通过极少量人类轨迹数据实现高效微调,从而快速低成本地适应新场景。它在长周期与灵巧任务中表现优异,包括需要双手协作与移动操作的任务,展现出稳健可靠的性能。这些能力得益于多阶段训练策略:包括与网络规模的视觉-语言数据共同训练、通过VR设备收集的人类轨迹高效微调,以及基于机器人轨迹的有效模仿学习。此外,我们推出了ByteMini——一款具出色灵活性与可靠性的双臂移动机器人,集成GR-3后可完成多样任务。大量真实世界实验表明,GR-3在多种挑战性任务上超越了当前最优基线方法π₀。我们希望GR-3能成为迈向日常生活中辅助人类的通用机器人的重要一步。
原文摘要 · Abstract (English)
We report our recent progress towards building generalist robot policies, the development of GR-3. GR-3 is a large-scale vision-language-action (VLA) model. It showcases exceptional capabilities in generalizing to novel objects, environments, and instructions involving abstract concepts. Furthermore, it can be efficiently fine-tuned with minimal human trajectory data, enabling rapid and cost-effective adaptation to new settings. GR-3 also excels in handling long-horizon and dexterous tasks, including those requiring bi-manual manipulation and mobile movement, showcasing robust and reliable performance. These capabilities are achieved through a multi-faceted training recipe that includes co-training with web-scale vision-language data, efficient fine-tuning from human trajectory data collected via VR devices, and effective imitation learning with robot trajectory data. In addition, we introduce ByteMini, a versatile bi-manual mobile robot designed with exceptional flexibility and reliability, capable of accomplishing a wide range of tasks when integrated with GR-3. Through extensive real-world experiments, we show GR-3 surpasses the state-of-the-art baseline method, $π_0$, on a wide variety of challenging tasks. We hope GR-3 can serve as a step towards building generalist robots capable of assisting humans in daily life.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。