arXiv:2410.06158cs.ROcs.CV2024-10被引 299

用3800万视频预训练,让机器人学会通用操作

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

  • 用3800万网络视频预训练,学习世界动态
  • 在100多个任务上平均成功率97.7%
  • 能泛化到新环境、新物体和新任务

我们提出GR-2,一种先进的通用机器人代理,可实现多样化且可泛化的机器人操作。GR-2首先在大量互联网视频上进行预训练,以捕捉世界的动态。这一大规模预训练涉及3800万段视频片段和超过500亿个词元,使GR-2在后续策略学习中具备跨多种机器人任务和环境的泛化能力。随后,GR-2通过机器人轨迹进行微调,用于视频生成和动作预测。其展现出卓越的多任务学习能力,在超过100个任务上达到平均97.7%的成功率。此外,GR-2在面对全新、未见过的场景(包括新背景、新环境、新物体和新任务)时表现出极强的泛化能力。值得注意的是,GR-2在模型规模增大时仍能有效扩展,展现出持续增长和应用的潜力。

原文摘要 · Abstract (English)

We present GR-2, a state-of-the-art generalist robot agent for versatile and generalizable robot manipulation. GR-2 is first pre-trained on a vast number of Internet videos to capture the dynamics of the world. This large-scale pre-training, involving 38 million video clips and over 50 billion tokens, equips GR-2 with the ability to generalize across a wide range of robotic tasks and environments during subsequent policy learning. Following this, GR-2 is fine-tuned for both video generation and action prediction using robot trajectories. It exhibits impressive multi-task learning capabilities, achieving an average success rate of 97.7% across more than 100 tasks. Moreover, GR-2 demonstrates exceptional generalization to new, previously unseen scenarios, including novel backgrounds, environments, objects, and tasks. Notably, GR-2 scales effectively with model size, underscoring its potential for continued growth and application. Project page: \url{https://gr2-manipulation.github.io}.

机器人操作视频预训练多任务学习泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。