让AI像设计师一样对话式编辑3D场景,精准定位且多视角一致。
DesignAgent3D: Interactive 3D Scene Editing via Designer-like Multimodal Reasoning

- 通过对话澄清模糊需求,分步规划、感知与执行编辑
- 在NeRF和3DGS上均实现高精度空间定位与多视角一致性
- 适合需要交互式、精确3D场景修改的研究者与设计师
文本引导的3D场景编辑为重构环境的修改提供了直观接口,但因自然语言设计请求语义不明确,需在杂乱3D场景中进行语义锚定,仍具挑战。现有方法通常将任务建模为单次提示下的条件生成,难以解决用户意图模糊问题,导致严重对象定位漂移、遮挡下跟踪失败以及多视角“贴纸效应”。为此,我们提出DesignAgent3D,一个交互式多模态智能体框架,将3D场景编辑重新定义为类设计师的“规划-感知-行动”范式。智能体首先通过与用户交互明确未指定的设计目标,然后在3D场景中精确定位目标对象或区域,最后施加受控视觉修改并保持场景一致性。编辑结果被集成到底层3D表示中,支持持久化和多视角一致的新视图渲染。在NeRF和3D Gaussian Splatting两种骨干网络上的大量实验表明,DesignAgent3D显著优于现有最优基线,在语义意图对齐、空间定位精度和多视角保真度方面均有卓越表现。
原文摘要 · Abstract (English)
Text guided 3D scene editing provides an intuitive interface for modifying reconstructed environments, but remains difficult because natural language design requests are often semantically underspecified and must be grounded in cluttered 3D scenes. Existing methods typically formulate the task as one-shot conditional generation from a single prompt, failing to resolve ambiguous user intents or achieve precise spatial grounding. Consequently, they suffer from severe object localization drift, tracking failure under occlusions, and the notorious multi-view "sticker effect." To overcome these limitations, we present DesignAgent3D, an interactive multimodal agentic framework that reformulates 3D scene editing as a designer-like Plan-Perceive-Act paradigm. The agent first plans by interacting with the user to clarify underspecified design goals, then perceives by grounding the intended edit to specific objects or regions in the 3D scene, and finally acts by applying controlled visual modifications while preserving scene consistency. The edits are further integrated into the underlying 3D representation, supporting persistent and multi-view consistent novel-view rendering. Extensive experiments across both NeRF and 3D Gaussian Splatting backbones demonstrate that DesignAgent3D significantly outperforms state-of-the-art baselines, delivering superior semantic intent alignment, impeccable spatial localization accuracy, and high-fidelity multi-view consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。