2D图像模型能当3D世界生成的智能体,让机器学会理解空间。
WorldAgents: Can Foundation Image Models be Agents for 3D World Models?
- 用多智能体架构让图像模型自主生成并验证3D视图。
- 在多个模型上实现连贯且一致的3D场景重建。
- 适合研究视觉语言模型与3D生成的学者参考。
鉴于2D基础图像模型生成高保真输出的卓越能力,我们探究一个根本问题:2D基础图像模型是否内在具备3D世界建模能力?为回答此问题,我们系统评估了多个前沿图像生成模型和视觉语言模型(VLMs)在3D世界合成任务上的表现。为挖掘并评测其潜在的隐式3D能力,我们提出一种智能体框架以促进3D世界生成。该方法采用多智能体架构:基于VLM的导演负责制定提示引导图像合成,生成器负责合成新视角图像,基于VLM的两步验证器则从2D图像与3D重建空间中评估并筛选生成帧。关键在于,我们的智能体方法实现了连贯且稳健的3D重建,生成的场景可支持新视角渲染。通过在多种基础模型上的广泛实验,我们证明2D模型确实蕴含对3D世界的理解。借助这一认知,本方法成功合成出广阔、真实且3D一致的世界。
原文摘要 · Abstract (English)
Given the remarkable ability of 2D foundation image models to generate high-fidelity outputs, we investigate a fundamental question: do 2D foundation image models inherently possess 3D world model capabilities? To answer this, we systematically evaluate multiple state-of-the-art image generation models and Vision-Language Models (VLMs) on the task of 3D world synthesis. To harness and benchmark their potential implicit 3D capability, we propose an agentic framing to facilitate 3D world generation. Our approach employs a multi-agent architecture: a VLM-based director that formulates prompts to guide image synthesis, a generator that synthesizes new image views, and a VLM-backed two-step verifier that evaluates and selectively curates generated frames from both 2D image and 3D reconstruction space. Crucially, we demonstrate that our agentic approach provides coherent and robust 3D reconstruction, producing output scenes that can be explored by rendering novel views. Through extensive experiments across various foundation models, we demonstrate that 2D models do indeed encapsulate a grasp of 3D worlds. By exploiting this understanding, our method successfully synthesizes expansive, realistic, and 3D-consistent worlds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。