让多模态智能体从用户指令端到端构建3D开放世界,挑战真实场景理解与编辑能力。
VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

- 提出统一框架VibeWorlding,支持意图推理、场景规划与多轮交互反馈
- 在6,828个反向合成查询上测试,最先进模型成功率仍不足60%
- 通过强化学习训练,开源模型可超越闭源前沿,30B版本表现最优
从用户查询构建可交互的3D开放世界至关重要。但现有方法多基于理想化、简单的查询进行评估,难以系统分析多模态智能体对用户意图的理解、3D工具使用及文本与视觉3D信息推理能力。为此,我们提出VibeWorlding:一个用于基准测试与训练“ vibe worlding agent”(多模态智能体)的统一框架——该智能体能自主推断用户意图、规划场景布局、调用3D工具,并在多轮人-机-环境交互中反思多模态反馈。为此,我们首先构建了VWE-BENCH:包含2,616个高质量3D资产、323个人工标注的种子3D世界及6,828个反向合成的多模态用户查询,分为带真值的验证查询与具严格评分标准的未验证查询。同时开发了VibeWorlding-Gym:一个联合多模态强化学习后训练框架,整合(1)统一沙盒环境(集成资产检索、编辑与图像渲染为MCP工具),以及(2)基于评分的验证器,结合物理可行性与意图满足度验证,支持公平模型评估与可扩展的多模态强化学习奖励服务。实验表明,当前前沿多模态大语言模型(MLLMs)远未解决该任务,即使GPT-5.5和Qwen3.8-Max成功率也低于60%,瓶颈在于精确3D世界编辑。进一步发现,强化学习训练可缓解此弱点,使开源模型甚至超越闭源前沿:我们的VibeWorlder-8B与前沿模型相当,旗舰版VibeWorlder-30B-A3B在所有评估模型中取得最佳整体Pass@1表现。
原文摘要 · Abstract (English)
Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it difficult to systematically analyze and compare how multimodal agents understand user intent, use 3D tools, and reason over textual and visual 3D world information. To this end, we propose VibeWorlding, a unified framework for benchmarking and training vibe worlding agents: a multimodal agent that can autonomously infer user intent, plan scene layout, invoke 3D tools, and reflect on the multimodal feedback in a multi-turn agent-environment interaction process. To achieve this, we first build VWE-BENCH, a benchmark of 2,616 high-quality 3D assets, 323 human-annotated seed 3D worlds, and 6,828 reverse-synthesized multimodal user queries, split into verified queries with ground-truth and unverified queries with carefully designed rubrics. Moreover, we develop VibeWorlding-Gym, a joint multimodal RL post-training framework that integrates (1) a sandbox environment unifying asset retrieval, editing, and image rendering as MCP tools, and (2) a rubric-based verifier that combines physical feasibility and intent fulfillment verification, supporting both fair model evaluation and scalable multimodal RL reward service. Our experiments show that current frontier MLLMs are far from solving the vibe worlding agent task, with even GPT-5.5 and Qwen3.8-Max reaching below 60% success rate, and trace the bottleneck to precise 3D world editing. We further find that RL training can ease this weakness and enable open-source MLLMs to even surpass closed-source frontiers: our VibeWorlder-8B is comparable to frontier MLLMs, while our flagship VibeWorlder-30B-A3B attains the best overall Pass@1 among all evaluated models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。