SIMA 2能理解复杂指令,在虚拟世界中自主学习并接近人类表现。
SIMA 2: A Generalist Embodied Agent for Virtual Worlds
- 基于Gemini模型,可理解语言与图像混合指令
- 在多种游戏环境中表现接近人类,且能泛化到新环境
- 能自动生成任务并自我改进,具备持续学习能力
我们提出SIMA 2,一种能在多种3D虚拟世界中理解与行动的通用体化智能体。基于Gemini基础模型,它实现了目标导向的主动交互,突破了以往仅支持简单语言指令的局限。相比SIMA 1,SIMA 2能进行高阶目标推理、与用户对话,并处理包含语言和图像的复杂指令。在多类游戏环境中,其性能显著缩小与人类的差距,且在未见过的环境中仍具强泛化能力,同时保持基础模型的核心推理能力。此外,通过Gemini生成任务与奖励,SIMA 2可在新环境中自主学习新技能。该工作验证了构建通用且持续进化智能体的可行性,为虚拟乃至物理世界中的智能体发展提供了路径。
原文摘要 · Abstract (English)
We introduce SIMA 2, a generalist embodied agent that understands and acts in a wide variety of 3D virtual worlds. Built upon a Gemini foundation model, SIMA 2 represents a significant step toward active, goal-directed interaction within an embodied environment. Unlike prior work (e.g., SIMA 1) limited to simple language commands, SIMA 2 acts as an interactive partner, capable of reasoning about high-level goals, conversing with the user, and handling complex instructions given through language and images. Across a diverse portfolio of games, SIMA 2 substantially closes the gap with human performance and demonstrates robust generalization to previously unseen environments, all while retaining the base model's core reasoning capabilities. Furthermore, we demonstrate a capacity for open-ended self-improvement: by leveraging Gemini to generate tasks and provide rewards, SIMA 2 can autonomously learn new skills from scratch in a new environment. This work validates a path toward creating versatile and continuously learning agents for both virtual and, eventually, physical worlds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。