让3D模型像2D一样用自然语言精准编辑,保持身份一致
DreamEdit3D: Personalization of Multi-View Diffusion Models for 3D Editing

- 分视角提取语义掩码,为每个物体学习独立的词嵌入
- 多视图一致性编辑效果优于现有方法,身份保留率高
- 适合需要个性化3D内容生成的设计师与开发者
尽管2D扩散模型在保持身份的个性化方面取得显著进展,但将这一能力扩展到3D资产仍面临多视角一致性与空间控制的挑战。受2D成果启发,我们提出一种新型文本引导3D编辑个性化方法,通过自然语言实现组合式、对象级控制。给定3D输入后,渲染正交视图并提取对象级分割掩码以分离语义组件。随后采用两阶段优化策略:多视图文本逆向生成结合注意力对齐,再对多视图扩散模型进行全量微调,学习各组件的独特词嵌入。推理时,这些解耦的词嵌入可无缝组合编辑提示,生成多视图一致图像,并进一步重建为高保真纹理3D网格。在多种编辑场景下的广泛评估表明,该方法成功将2D个性化的灵活性迁移至3D,相较于现有基线,在编辑忠实度与身份保留方面达到最先进水平。
原文摘要 · Abstract (English)
While 2D diffusion models have achieved remarkable success in identity-preserving personalization, extending this capability to 3D assets remains a significant challenge due to the complexities of multi-view consistency and spatial control. Inspired by these 2D advancements, we present a novel personalization method for text-guided 3D editing that enables compositional, object-level control through natural language. Given a 3D input, we render orthogonal views and extract object-level segmentation masks to isolate semantic components. We then learn distinct token embeddings for each component through a tailored two-phase optimization strategy: multi-view textual inversion with attention alignment, followed by full fine-tuning of multi-view diffusion model. During inference, these disentangled tokens seamlessly compose with editing prompts to generate multi-view consistent images, which are subsequently lifted into high-fidelity textured 3D meshes. Extensive evaluations across diverse editing scenarios demonstrate that our method successfully transfers the flexibility of 2D personalization to 3D, achieving state-of-the-art edit faithfulness and identity preservation compared to existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。