arXiv:2503.19452cs.CV2025-03被引 7

仅用5张照片重建复杂户外场景,还能处理遮挡和光照变化。

SparseGS-W: Sparse-View 3D Gaussian Splatting in the Wild with Generative Priors

  • 基于3D高斯溅射,结合生成先验补足稀疏视图信息。
  • 在仅有5张输入图像时仍保持高质量新视角合成,无明显伪影。
  • 适合野外稀疏拍摄场景,尤其适用于光照一致的重建任务。

从非受限的野外图像中合成大规模场景的新视角,是计算机视觉中的重要但具挑战性的任务。现有方法依赖约1000张密集训练图像,通过隐式神经网络优化每张图像的外观和瞬时遮挡,但在稀疏输入下表现不佳,产生明显伪影。为此,我们提出SparseGS-W,一种基于3D高斯溅射的新框架,仅需5张训练图像即可重建复杂户外场景,并有效处理遮挡与外观变化。该方法利用几何先验和受控扩散先验弥补极稀疏输入带来的多视角信息缺失。具体而言,提出即插即用的受控新视角增强模块,在高斯优化过程中迭代提升渲染质量;同时设计遮挡处理模块,借助受控扩散模型的高质量修复能力灵活消除遮挡。两个模块均可从任意用户提供参考图像中提取外观特征,实现光照一致性建模。在PhotoTourism和Tanks and Temples数据集上的大量实验表明,SparseGS-W不仅在全参考指标上达到最先进水平,还在常用无参考指标(如FID、ClipIQA、MUSIQ)上表现优异。

原文摘要 · Abstract (English)

Synthesizing novel views of large-scale scenes from unconstrained in-the-wild images is an important but challenging task in computer vision. Existing methods, which optimize per-image appearance and transient occlusion through implicit neural networks from dense training views (approximately 1000 images), struggle to perform effectively under sparse input conditions, resulting in noticeable artifacts. To this end, we propose SparseGS-W, a novel framework based on 3D Gaussian Splatting that enables the reconstruction of complex outdoor scenes and handles occlusions and appearance changes with as few as five training images. We leverage geometric priors and constrained diffusion priors to compensate for the lack of multi-view information from extremely sparse input. Specifically, we propose a plug-and-play Constrained Novel-View Enhancement module to iteratively improve the quality of rendered novel views during the Gaussian optimization process. Furthermore, we propose an Occlusion Handling module, which flexibly removes occlusions utilizing the inherent high-quality inpainting capability of constrained diffusion priors. Both modules are capable of extracting appearance features from any user-provided reference image, enabling flexible modeling of illumination-consistent scenes. Extensive experiments on the PhotoTourism and Tanks and Temples datasets demonstrate that SparseGS-W achieves state-of-the-art performance not only in full-reference metrics, but also in commonly used non-reference metrics such as FID, ClipIQA, and MUSIQ.

3D重建高斯溅射稀疏视图生成先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。