通过对应点一致性指导,实现多视角图像编辑的3D几何一致
3D-Consistent Multi-View Editing by Correspondence Guidance
- 基于对应点视觉相似性设计一致性损失,引导去噪过程
- 无需训练即可提升多视图编辑的3D一致性,支持稀疏与密集编辑
- 适用于NeRF、高斯点云等3D表示,适配各类文本编辑模型
扩散模型和流模型的进步显著提升了文本驱动的图像编辑能力,但独立编辑各视角图像常导致几何与光照不一致,尤其在NeRF或高斯点云等3D表示中问题突出。本文提出一种无需训练的一致性引导框架,在编辑过程中强制多视角一致性。核心思想是:对应点在编辑后应保持视觉相似。为此引入一致性损失,引导去噪过程趋向一致编辑。该框架灵活,可与多种图像编辑方法结合,支持稀疏与密集多视角编辑设置。实验表明,本方法显著优于现有方法,在3D一致性方面表现更优。同时,更高的一致性使高斯点云编辑能保留清晰细节,并忠实于用户指定的文本提示。视频结果见项目主页:https://3d-consistent-editing.github.io/
原文摘要 · Abstract (English)
Recent advancements in diffusion and flow models have greatly improved text-based image editing, yet methods that edit images independently often produce geometrically and photometrically inconsistent results across different views of the same scene. Such inconsistencies are particularly problematic for editing of 3D representations such as NeRFs or Gaussian splat models. We propose a training-free guidance framework that enforces multi-view consistency during the image editing process. The key idea is that corresponding points should look similar after editing. To achieve this, we introduce a consistency loss that guides the denoising process toward coherent edits. The framework is flexible and can be combined with widely varying image editing methods, supporting both dense and sparse multi-view editing setups. Experimental results show that our approach significantly improves 3D consistency compared to existing multi-view editing methods. We also show that this increased consistency enables high-quality Gaussian splat editing with sharp details and strong fidelity to user-specified text prompts. Please refer to our project page for video results: https://3d-consistent-editing.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。