arXiv:2503.11601cs.CV2025-03被引 3

提升文本控制3D高斯点云编辑的视觉质量和多视角一致性

Advancing 3D Gaussian Splatting Editing with Complementary and Consensus Information

  • 通过互补信息互学习网络优化深度图估计
  • 利用小波共识注意力实现多视角结果一致
  • 适合需要高质量3D场景编辑的研究者

我们提出一种新框架,用于提升文本引导的3D高斯点云(3DGS)编辑在视觉保真度和一致性方面的表现。现有方法存在两大挑战:在复杂相机位姿下多视角几何重建不一致,以及图像操作中深度信息利用不足,导致纹理过度、物体边界模糊。为此,我们引入:1)互补信息互学习网络,提升3DGS的深度图估计精度,实现基于深度条件的3D编辑并保持几何结构;2)小波共识注意力机制,在扩散去噪过程中对齐潜在编码,确保编辑结果的多视角一致性。大量实验表明,该方法在渲染质量和视角一致性上均优于当前最优方法,验证了其在文本引导3D场景编辑中的有效性。

原文摘要 · Abstract (English)

We present a novel framework for enhancing the visual fidelity and consistency of text-guided 3D Gaussian Splatting (3DGS) editing. Existing editing approaches face two critical challenges: inconsistent geometric reconstructions across multiple viewpoints, particularly in challenging camera positions, and ineffective utilization of depth information during image manipulation, resulting in over-texture artifacts and degraded object boundaries. To address these limitations, we introduce: 1) A complementary information mutual learning network that enhances depth map estimation from 3DGS, enabling precise depth-conditioned 3D editing while preserving geometric structures. 2) A wavelet consensus attention mechanism that effectively aligns latent codes during the diffusion denoising process, ensuring multi-view consistency in the edited results. Through extensive experimentation, our method demonstrates superior performance in rendering quality and view consistency compared to state-of-the-art approaches. The results validate our framework as an effective solution for text-guided editing of 3D scenes.

3D高斯图像编辑多视角一致扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。