arXiv:2503.14734cs.ROcs.AI2025-03被引 1.3k

开源通用人形机器人基础模型,能听懂指令并实时生成流畅动作。

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

论文配图:GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
图 1 · 摘自论文原文
  • 双系统架构:视觉语言理解+扩散动作生成,端到端联合训练。
  • 在多个机器人形态上超越现有模仿学习方法,仿真测试表现优异。
  • 高数据效率,可在真实人形机器人上完成语言控制的双手操作。

通用机器人需要多功能身体和智能大脑。近年来,人形机器人作为实现通用自主性的硬件平台展现出巨大潜力。一个在海量多样化数据上训练的机器人基础模型,对于使机器人能够推理新情境、稳健应对现实世界变化并快速学习新任务至关重要。为此,我们提出GR00T N1,一个面向人形机器人的开源基础模型。GR00T N1是一种视觉-语言-动作(VLA)模型,采用双系统架构:视觉语言模块(系统2)通过视觉和语言指令理解环境;后续的扩散变换器模块(系统1)实时生成流畅运动动作。两个模块紧密耦合,端到端联合训练。我们使用真实机器人轨迹、人类视频和合成数据的异构混合数据集对GR00T N1进行训练。实验表明,该通用机器人模型在多个机器人形态的标准仿真基准上优于当前最优的模仿学习基线。此外,我们在Fourier GR-1人形机器人上部署该模型,实现了语言引导下的双臂操作任务,表现出色且具有高数据效率。

原文摘要 · Abstract (English)

General-purpose robots need a versatile body and an intelligent mind. Recent advancements in humanoid robots have shown great promise as a hardware platform for building generalist autonomy in the human world. A robot foundation model, trained on massive and diverse data sources, is essential for enabling the robots to reason about novel situations, robustly handle real-world variability, and rapidly learn new tasks. To this end, we introduce GR00T N1, an open foundation model for humanoid robots. GR00T N1 is a Vision-Language-Action (VLA) model with a dual-system architecture. The vision-language module (System 2) interprets the environment through vision and language instructions. The subsequent diffusion transformer module (System 1) generates fluid motor actions in real time. Both modules are tightly coupled and jointly trained end-to-end. We train GR00T N1 with a heterogeneous mixture of real-robot trajectories, human videos, and synthetically generated datasets. We show that our generalist robot model GR00T N1 outperforms the state-of-the-art imitation learning baselines on standard simulation benchmarks across multiple robot embodiments. Furthermore, we deploy our model on the Fourier GR-1 humanoid robot for language-conditioned bimanual manipulation tasks, achieving strong performance with high data efficiency.

人形机器人基础模型视觉语言动作仿生控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。