让3D编辑能理解语义,用多图+文本精准修改物体局部。
ES3D: Embedding Semantics into 3D Space for Component-Aware Editing

- 将多视角语义特征投影到3D体素空间,构建可检索的语义嵌入。
- 支持基于图像或文本的组件定位,编辑时保持其余部分不变。
- 适合需要精细控制3D模型局部修改的设计师与开发者。
现有3D编辑方法在可控性上取得进展,但仍存在局限:多数依赖文本驱动,难以表达用户意图的细粒度视觉变化;且常需手动提供3D掩码,或对不应修改区域造成干扰。这些限制源于缺乏细粒度语义理解,导致模型难以精准定位或修改特定3D组件。本文提出ES3D框架,将语义直接嵌入3D空间,实现基于多张局部参考图像和可选文本查询的组件感知检索与编辑。首先,通过将多视图语义特征投影至资产的体素空间,构建3D语义嵌入;接着,通过计算该嵌入与图像或文本查询的语义嵌入间的特征相似性,完成3D组件检索;最后,利用预训练3D生成模型结合修复机制,在用户提供的图像引导下修改目标组件,同时保持其余部分不变。实验表明,ES3D生成的编辑结果在几何与语义上均具一致性,支持鲁棒的基于图像与文本辅助的3D编辑控制。
原文摘要 · Abstract (English)
Existing 3D editing methods have made notable progress in controllability, yet they remain limited in several important ways. Most approaches rely on text-driven editing, which struggles to express fine-grained visual changes intended by the user. Moreover, many methods require manually supplied 3D masks or introduce unintended changes to regions that should remain untouched. These limitations largely arise from the absence of fine-grained semantic understanding, making it difficult for existing models to retrieve or modify specific 3D components. We introduce ES3D, a framework that embeds semantics directly into 3D space, enabling component-aware retrieval and editing of a 3D asset conditioned on multiple local reference images and optional text queries. We first construct a 3D semantic embedding by projecting multi-view semantic features into the voxelized space of the asset. We then perform 3D component retrieval by computing feature similarity between the 3D semantic embedding and the semantic embeddings of image or text queries. For editing, we employ a pretrained 3D generative model with an inpainting mechanism to modify the retrieved components guided by user-provided images while preserving the rest of the asset. Overall, ES3D is a 3D editing framework that retrieves editable regions based on semantic cues and uses multiple images as conditions. Extensive experiments demonstrate that ES3D produces geometrically consistent and semantically coherent edits, enabling robust image-based and text-assisted control for 3D editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。