无需空间提示,用自然语言在3D场景中灵活插入物体。
FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors
- 分离生成与定位,通过多模态模型解析语义
- 无需2D掩码或3D框,实现精准且一致的物体插入
- 适合需要快速编辑3D场景的设计师和开发者
基于文本的3D场景物体插入是一项新兴任务,支持通过自然语言进行直观编辑。然而,现有基于2D编辑的方法通常依赖于2D掩码或3D边界框等空间先验,难以保证插入物体的一致性,限制了其在真实场景中的灵活性与可扩展性。本文提出FreeInsert,一种新颖框架,利用基础模型(包括MLLMs、LGMs和扩散模型)将物体生成与空间放置解耦,实现无需空间先验的无监督、灵活物体插入。FreeInsert首先通过MLLM解析用户指令中的结构化语义,包括物体类型、空间关系和附着区域;这些语义同时指导物体重建以保证3D一致性,并学习其自由度。借助MLLM的空间推理能力初始化物体姿态与尺度。随后,分层的空间感知优化阶段融合空间语义与MLLM推断的先验,提升放置精度。最后,使用插入物体图像增强外观,提高视觉保真度。实验表明,FreeInsert在不依赖空间先验的前提下,实现了语义一致、空间精确、视觉真实的3D插入,提供友好且灵活的编辑体验。
原文摘要 · Abstract (English)
Text-driven object insertion in 3D scenes is an emerging task that enables intuitive scene editing through natural language. However, existing 2D editing-based methods often rely on spatial priors such as 2D masks or 3D bounding boxes, and they struggle to ensure consistency of the inserted object. These limitations hinder flexibility and scalability in real-world applications. In this paper, we propose FreeInsert, a novel framework that leverages foundation models including MLLMs, LGMs, and diffusion models to disentangle object generation from spatial placement. This enables unsupervised and flexible object insertion in 3D scenes without spatial priors. FreeInsert starts with an MLLM-based parser that extracts structured semantics, including object types, spatial relationships, and attachment regions, from user instructions. These semantics guide both the reconstruction of the inserted object for 3D consistency and the learning of its degrees of freedom. We leverage the spatial reasoning capabilities of MLLMs to initialize object pose and scale. A hierarchical, spatially aware refinement stage further integrates spatial semantics and MLLM-inferred priors to enhance placement. Finally, the appearance of the object is improved using the inserted-object image to enhance visual fidelity. Experimental results demonstrate that FreeInsert achieves semantically coherent, spatially precise, and visually realistic 3D insertions without relying on spatial priors, offering a user-friendly and flexible editing experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。