Gemini Robotics让机器人直接理解并执行复杂指令,实现物理世界通用智能控制。
Gemini Robotics: Bringing AI into the Physical World
- 基于Gemini 2.0构建视觉-语言-动作统一模型,直接驱动机器人完成任务。
- 可处理未见过的环境与开放词汇指令,支持少至100次演示的新任务学习。
- 增强空间时间理解能力,适用于多视角对齐、3D目标检测等机器人关键场景。
大型多模态模型在数字领域展现出强大的通用能力,但将其应用于机器人等物理代理仍面临巨大挑战。本文提出Gemini Robotics,一个专为机器人设计的先进视觉-语言-动作(VLA)通用模型,能够直接控制机器人执行复杂操作。该模型可实现流畅、实时响应的动作,适应不同物体类型与位置变化,应对未知环境,并遵循多样化的开放式词汇指令。通过进一步微调,其可拓展至长时程高灵巧任务求解、仅需100次示范即学会新短时任务,甚至适配全新机器人形态。这一能力依托于我们提出的Gemini Robotics-ER(具身推理)模型,该模型将Gemini的多模态推理能力延伸至物理世界,增强了空间与时间理解,支持物体检测、指指点位、轨迹与抓取预测、多视角对应关系及3D边界框预测。实验验证了该组合在多种机器人应用中的有效性。同时,我们探讨并回应了此类具身基础模型带来的安全问题。Gemini Robotics系列标志着迈向通用机器人、实现人工智能在物理世界潜力的重要一步。
原文摘要 · Abstract (English)
Recent advancements in large multimodal models have led to the emergence of remarkable generalist capabilities in digital domains, yet their translation to physical agents such as robots remains a significant challenge. This report introduces a new family of AI models purposefully designed for robotics and built upon the foundation of Gemini 2.0. We present Gemini Robotics, an advanced Vision-Language-Action (VLA) generalist model capable of directly controlling robots. Gemini Robotics executes smooth and reactive movements to tackle a wide range of complex manipulation tasks while also being robust to variations in object types and positions, handling unseen environments as well as following diverse, open vocabulary instructions. We show that with additional fine-tuning, Gemini Robotics can be specialized to new capabilities including solving long-horizon, highly dexterous tasks, learning new short-horizon tasks from as few as 100 demonstrations and adapting to completely novel robot embodiments. This is made possible because Gemini Robotics builds on top of the Gemini Robotics-ER model, the second model we introduce in this work. Gemini Robotics-ER (Embodied Reasoning) extends Gemini's multimodal reasoning capabilities into the physical world, with enhanced spatial and temporal understanding. This enables capabilities relevant to robotics including object detection, pointing, trajectory and grasp prediction, as well as multi-view correspondence and 3D bounding box predictions. We show how this novel combination can support a variety of robotics applications. We also discuss and address important safety considerations related to this new class of robotics foundation models. The Gemini Robotics family marks a substantial step towards developing general-purpose robots that realizes AI's potential in the physical world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。