打造逼真可交互的虚拟世界,让大模型智能体真实练就生存技能。
SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds
- 基于虚幻引擎5构建开放世界,支持语言生成动态环境与真实物理社交规则。
- 部署GPT-4o等前沿大模型完成多智能体长周期配送任务,揭示不同模型推理差异。
- 开源平台适合研究真实世界智能体,尤其关注具身智能与社会协作方向。
尽管大型语言模型(LLM)和视觉语言模型(VLM)在数学、编程和计算机使用方面进展迅速,但其在复杂物理与社会环境中的应用仍面临挑战。要使智能体在现实世界中自主生存与发展(如独立创收或经营业务),需在多样化具身场景中进行大规模交互、推理、训练与评估。然而,现有世界模拟器存在局限:通常依赖有限的手工设计环境,模拟简化版游戏式物理与社会规则,且缺乏对LLM/VLM智能体的原生支持。我们提出SimWorld,一个基于虚幻引擎5构建的新一代模拟器,旨在为开发与评估LLM/VLM智能体提供丰富、逼真的现实世界级环境。SimWorld具备三大核心能力:(1) 真实、开放式的世界仿真,包含精准的物理与社会动态,以及语言驱动的程序化环境生成;(2) 面向LLM/VLM智能体的丰富接口,支持多模态输入与多层次抽象的开放式词汇动作;(3) 易于用户自定义的多样化物理与社会推理场景。我们通过部署前沿大模型(如GPT-4o、Gemini-2.5-Flash、Claude-3.5与DeepSeek-Prover-V2)执行涉及策略合作与竞争的长时程多智能体配送任务,揭示了各模型在推理模式与能力边界上的显著差异。我们已开源SimWorld,期望其成为推动跨学科真实世界智能体发展的基础平台:https://simworld.org。
原文摘要 · Abstract (English)
While LLM/VLM-powered AI agents have advanced rapidly in math, coding, and computer use, their applications in complex physical and social environments remain challenging. Building agents that can survive and thrive in the real world (for example, by autonomously earning income or running a business) requires massive-scale interaction, reasoning, training, and evaluation across diverse embodied scenarios. However, existing world simulators for such development fall short: they often rely on limited hand-crafted environments, simulate simplified game-like physics and social rules, and lack native support for LLM/VLM agents. We introduce SimWorld, a new simulator built on Unreal Engine 5, designed for developing and evaluating LLM/VLM agents in rich, real-world-like settings. SimWorld offers three core capabilities: (1) realistic, open-ended world simulation, including accurate physical and social dynamics and language-driven procedural environment generation; (2) a rich interface for LLM/VLM agents, with multimodal world inputs and open-vocabulary actions at varying levels of abstraction; and (3) diverse and extensible physical and social reasoning scenarios that are easily customizable by users. We demonstrate SimWorld by deploying frontier LLM agents (e.g., GPT-4o, Gemini-2.5-Flash, Claude-3.5, and DeepSeek-Prover-V2) on long-horizon multi-agent delivery tasks involving strategic cooperation and competition. The results reveal distinct reasoning patterns and limitations across models. We open-source SimWorld and hope it becomes a foundational platform for advancing real-world agent intelligence across disciplines: https://simworld.org.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。