arXiv:2603.01603cs.CV2026-03

解决稀疏视角下3D高斯点云的瞬时干扰物问题

Sparse View Distractor-Free Gaussian Splatting

  • 用几何模型预估相机参数并生成稠密初始点
  • 通过注意力图实现语义实体精准匹配
  • 结合视觉语言模型保留大静态区域,适合稀疏场景重建

3D高斯点云(3DGS)在静态环境中实现了高效训练与快速新视角生成。针对动态物体带来的挑战,无干扰物3DGS方法已取得良好效果,但当输入图像稀疏时性能显著下降。主要原因在于依赖颜色残差启发式指导训练,而稀疏观测下该策略不可靠。本文提出一种新框架,在稀疏视角下增强无干扰物3DGS的表现,通过引入丰富先验信息:首先利用几何基础模型VGGT估计相机参数并生成稠密初始3D点;其次利用VGGT的注意力图实现高效准确的语义实体匹配;此外,借助视觉-语言模型(VLMs)进一步识别并保留场景中的大范围静态区域。我们还展示了这些先验如何无缝集成至现有无干扰物3DGS方法中。大量实验验证了本方法在缓解稀疏视角3DGS训练中瞬时干扰物方面的有效性与鲁棒性。

原文摘要 · Abstract (English)

3D Gaussian Splatting (3DGS) enables efficient training and fast novel view synthesis in static environments. To address challenges posed by transient objects, distractor-free 3DGS methods have emerged and shown promising results when dense image captures are available. However, their performance degrades significantly under sparse input conditions. This limitation primarily stems from the reliance on the color residual heuristics to guide the training, which becomes unreliable with limited observations. In this work, we propose a framework to enhance distractor-free 3DGS under sparse-view conditions by incorporating rich prior information. Specifically, we first adopt the geometry foundation model VGGT to estimate camera parameters and generate a dense set of initial 3D points. Then, we harness the attention maps from VGGT for efficient and accurate semantic entity matching. Additionally, we utilize Vision-Language Models (VLMs) to further identify and preserve the large static regions in the scene. We also demonstrate how these priors can be seamlessly integrated into existing distractor-free 3DGS methods. Extensive experiments confirm the effectiveness and robustness of our approach in mitigating transient distractors for sparse-view 3DGS training.

3D重建高斯点云稀疏视角视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。