arXiv:2509.20414cs.GRcs.CV2025-09NeurIPS被引 55

用智能代理迭代生成更真实多样的3D室内场景

SceneWeaver: All-in-One 3D Scene Synthesis with an Extensible and Self-Reflective Agent

  • 通过可扩展工具链与自我反思机制,动态选择生成策略
  • 在物理合理性、视觉真实性和指令对齐上超越现有方法
  • 适合需要复杂指令响应的通用3D环境构建任务

随着具身AI的发展,室内场景合成的重要性日益凸显,要求3D环境不仅视觉逼真,还需物理合理且功能多样。现有方法虽提升了视觉质量,但通常局限于固定场景类别,缺乏物体级细节与物理一致性,难以应对复杂用户指令。本文提出SceneWeaver,一个基于语言模型规划器的反射式智能体框架,通过工具驱动的迭代优化统一多种合成范式。该框架利用自评估机制,在物理合理性、视觉真实性和语义对齐性指导下,从数据驱动生成模型到视觉与大模型结合的方法中选择合适工具,实现‘思考-执行-反思’闭环。实验表明,SceneWeaver在常见及开放词汇房间类型上均优于已有方法,尤其在复杂指令下的泛化能力显著提升,推动通用3D环境生成发展。

原文摘要 · Abstract (English)

Indoor scene synthesis has become increasingly important with the rise of Embodied AI, which requires 3D environments that are not only visually realistic but also physically plausible and functionally diverse. While recent approaches have advanced visual fidelity, they often remain constrained to fixed scene categories, lack sufficient object-level detail and physical consistency, and struggle to align with complex user instructions. In this work, we present SceneWeaver, a reflective agentic framework that unifies diverse scene synthesis paradigms through tool-based iterative refinement. At its core, SceneWeaver employs a language model-based planner to select from a suite of extensible scene generation tools, ranging from data-driven generative models to visual- and LLM-based methods, guided by self-evaluation of physical plausibility, visual realism, and semantic alignment with user input. This closed-loop reason-act-reflect design enables the agent to identify semantic inconsistencies, invoke targeted tools, and update the environment over successive iterations. Extensive experiments on both common and open-vocabulary room types demonstrate that SceneWeaver not only outperforms prior methods on physical, visual, and semantic metrics, but also generalizes effectively to complex scenes with diverse instructions, marking a step toward general-purpose 3D environment generation. Project website: https://scene-weaver.github.io/.

3D生成智能体场景合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。