arXiv:2505.20469cs.CVcs.AI2025-05ICCV被引 8

解决3D高斯点云语义不一致问题,提升跨视角语义一致性。

CCL-LGS: Contrastive Codebook Learning for 3D Language Gaussian Splatting

  • 用零样本追踪对齐SAM生成的2D掩码并识别类别。
  • 通过对比码本学习增强类内紧凑性与类间区分性。
  • 适合需要高精度3D语义重建的机器人与自动驾驶场景。

近期3D重建技术和视觉语言模型的进步推动了3D语义理解的发展,这对机器人、自动驾驶及虚拟/增强现实至关重要。然而,依赖2D先验的方法易受遮挡、图像模糊和视图相关性变化的影响,导致跨视角语义不一致,经投影监督传播后会降低3D高斯语义场质量并引入渲染伪影。为此,我们提出CCL-LGS框架,通过整合多视角语义线索实现视图一致的语义监督。首先,利用零样本追踪对齐一组SAM生成的2D掩码,并可靠识别其对应类别;其次,使用CLIP提取跨视角鲁棒的语义编码;最后,对比码本学习(CCL)模块通过强制类内紧凑性和类间差异性,提炼出判别性语义特征。相比直接将CLIP应用于不完美掩码的方法,本框架显式解决语义冲突同时保持类别可区分性。大量实验表明,CCL-LGS优于先前最先进方法。

原文摘要 · Abstract (English)

Recent advances in 3D reconstruction techniques and vision-language models have fueled significant progress in 3D semantic understanding, a capability critical to robotics, autonomous driving, and virtual/augmented reality. However, methods that rely on 2D priors are prone to a critical challenge: cross-view semantic inconsistencies induced by occlusion, image blur, and view-dependent variations. These inconsistencies, when propagated via projection supervision, deteriorate the quality of 3D Gaussian semantic fields and introduce artifacts in the rendered outputs. To mitigate this limitation, we propose CCL-LGS, a novel framework that enforces view-consistent semantic supervision by integrating multi-view semantic cues. Specifically, our approach first employs a zero-shot tracker to align a set of SAM-generated 2D masks and reliably identify their corresponding categories. Next, we utilize CLIP to extract robust semantic encodings across views. Finally, our Contrastive Codebook Learning (CCL) module distills discriminative semantic features by enforcing intra-class compactness and inter-class distinctiveness. In contrast to previous methods that directly apply CLIP to imperfect masks, our framework explicitly resolves semantic conflicts while preserving category discriminability. Extensive experiments demonstrate that CCL-LGS outperforms previous state-of-the-art methods. Our project page is available at https://epsilontl.github.io/CCL-LGS/.

3D语义重建高斯泼溅视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。