用对应约束注意力实现文本驱动3D编辑的多视角一致性
CoreEditor: Correspondence-constrained Diffusion for Consistent 3D Editing
- 引入对应约束注意力机制,确保跨视角像素在去噪中保持一致
- 生成细节更锐利、一致性显著优于现有方法的3D编辑结果
- 支持用户从多个候选结果中选择,提升交互灵活性
文本驱动的3D编辑旨在根据文本描述修改3D场景,现有方法通常将预训练的2D图像编辑器适配到多视图输入。然而,由于缺乏对多视图信息交换的显式控制,这些方法常导致跨视角不一致,造成编辑不足和细节模糊。我们提出CoreEditor,一种新型的一致性文本到3D编辑框架。其核心创新是对应约束注意力机制,强制在扩散去噪过程中,预期保持一致的像素间进行精确交互。除几何对齐外,还结合去噪过程中估计的语义相似性,实现更可靠的对应建模与鲁棒的多视图编辑。此外,设计了选择性编辑流程,允许用户从多个候选结果中挑选偏好输出,提升灵活性与用户控制力。大量实验表明,CoreEditor生成的编辑结果质量高、3D一致性强,显著优于先前方法。
原文摘要 · Abstract (English)
Text-driven 3D editing seeks to modify 3D scenes according to textual descriptions, and most existing approaches tackle this by adapting pre-trained 2D image editors to multi-view inputs. However, without explicit control over multi-view information exchange, they often fail to maintain cross-view consistency, leading to insufficient edits and blurry details. We introduce CoreEditor, a novel framework for consistent text-to-3D editing. The key innovation is a correspondence-constrained attention mechanism that enforces precise interactions between pixels expected to remain consistent throughout the diffusion denoising process. Beyond relying solely on geometric alignment, we further incorporate semantic similarity estimated during denoising, enabling more reliable correspondence modeling and robust multi-view editing. In addition, we design a selective editing pipeline that allows users to choose preferred results from multiple candidates, offering greater flexibility and user control. Extensive experiments show that CoreEditor produces high-quality, 3D-consistent edits with sharper details, significantly outperforming prior methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。