arXiv:2504.17815cs.CV2025-04被引 1

用视觉不确定性指导3D高斯补全,让缺失物体自然融入场景。

Visibility-Uncertainty-guided 3D Gaussian Inpainting via Scene Conceptional Learning

  • 基于多视角可见性不确定性,动态融合互补视觉线索
  • 在SPIn-NeRF和水下数据集上实现无伪影的高质量补全
  • 支持静态与动态目标补全,适合复杂场景重建

3D高斯点云(3DGS)已成为高效的新视角合成3D表示方法。本文将3DGS拓展至补全任务,即替换场景中被遮挡物体并使其与周围自然融合。不同于2D图像补全,3D高斯补全需有效利用多视图间的互补视觉与语义信息,因某视角遮挡区域可能在其他视角可见。为此,我们提出一种方法,通过测量3D点在不同输入视图中的可见性不确定性,引导补全过程充分利用互补视觉线索。同时,利用不确定性学习不含遮挡物的场景语义概念,并基于该概念使用扩散模型在输入图像中填充被遮挡物体。最终,我们构建了新型3D补全框架VISTA,集成可见性不确定性引导的3DGI与场景概念学习。VISTA生成高质量3DGS模型,可合成无伪影且自然的新增视角。此外,本方法还可处理随时间变化的动态干扰物,在多种场景重建中展现更强泛化能力。我们在两个挑战性数据集上验证性能:包含10个多样静态场景的SPIn-NeRF数据集,以及基于UTB180构建的水下3D补全数据集,其中以快速游动的鱼为补全目标。

原文摘要 · Abstract (English)

3D Gaussian Splatting (3DGS) has emerged as a powerful and efficient 3D representation for novel view synthesis. This paper extends 3DGS capabilities to inpainting, where masked objects in a scene are replaced with new contents that blend seamlessly with the surroundings. Unlike 2D image inpainting, 3D Gaussian inpainting (3DGI) is challenging in effectively leveraging complementary visual and semantic cues from multiple input views, as occluded areas in one view may be visible in others. To address this, we propose a method that measures the visibility uncertainties of 3D points across different input views and uses them to guide 3DGI in utilizing complementary visual cues. We also employ uncertainties to learn a semantic concept of scene without the masked object and use a diffusion model to fill masked objects in input images based on the learned concept. Finally, we build a novel 3DGI framework, VISTA, by integrating VISibility-uncerTainty-guided 3DGI with scene conceptuAl learning. VISTA generates high-quality 3DGS models capable of synthesizing artifact-free and naturally inpainted novel views. Furthermore, our approach extends to handling dynamic distractors arising from temporal object changes, enhancing its versatility in diverse scene reconstruction scenarios. We demonstrate the superior performance of our method over state-of-the-art techniques using two challenging datasets: the SPIn-NeRF dataset, featuring 10 diverse static 3D inpainting scenes, and an underwater 3D inpainting dataset derived from UTB180, including fast-moving fish as inpainting targets.

3D补全高斯溅射扩散模型多视图融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。