用3D高斯点云引导2D图像修复,支持任意形状掩码和物体插入。
CoIn: Comprehensive 2D-3D Inpainting with Gaussian Splatting Guidance

- 通过扩散模型生成初始修复图,支持任意形状掩码。
- 利用参考视图自适应高斯点云重建3D场景,实现多视角一致性。
- 双向信息流融合,可同时完成去遮挡与物体插入任务。
3D场景修复对重建被遮挡或视角受限区域至关重要。现有方法虽利用高斯点云(GS)实现高效3D编辑,但通常依赖精确的多视角分割掩码,且仅适用于物体移除任务。本文提出CoIn,一种通过多阶段一致性流程连接2D修复模型与3D GS的新框架。首先使用扩散模型生成初始修复图像,支持任意形状掩码及物体插入等多样化任务。随后引入特征注意力引导的参考自适应高斯点云(Reference Adaptive GS),基于参考视图重构粗略3D场景(2D→3D)。该3D表示通过基于高斯点云的参考特征扭曲,为扩散过程提供几何引导,确保多视角一致性(3D→2D)。最后,纹理增强判别器对3D场景进行细化,提升光度真实感(2D→3D)。实验表明,CoIn通过双向信息流实现最优性能,在物体移除与插入任务中均表现优异,且支持灵活掩码输入。
原文摘要 · Abstract (English)
3D scene inpainting is essential for reconstructing areas corrupted by occlusions or limited viewpoints. While recent methods leverage Gaussian Splatting (GS) for efficient 3D editing, they often depend on precise multi-view segmentation masks and are inherently constrained to object removal tasks. We propose CoIn, a novel framework that bridges 2D inpainting models and 3DGS through a multi-stage consistency pipeline. Our approach first generates initial inpainted images using a diffusion model, enabling the use of arbitrary-shaped masks and diverse tasks like object insertion. We then introduce Reference Adaptive GS with Feature Attention to reconstruct a coarse 3D scene by adaptively weighing towards a reference view (2D -> 3D). This 3D representation provides geometric guidance to the diffusion process via GS-based Reference Feature Warping, ensuring multi-view consistency (3D -> 2D). Finally, a Texture-Enhancing Discriminator refines the 3D scene to achieve high photometric realism (2D -> 3D). Experiments show that CoIn, effectively leveraging bidirectional information flow, achieves state-of-the-art performance and effectively handles both object removal and object insertion with flexible mask input.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。