arXiv:2503.15712cs.CV2025-03

用几何先验提升CLIP在3D场景分割中的零样本能力

SPNeRF: Open Vocabulary 3D Neural Scene Segmentation with Superpoints

  • 通过超点几何结构增强NeRF训练,生成更精准的特征
  • 在不依赖额外模型下,性能超越原始LERF方法
  • 适合研究3D零样本分割与视觉语言模型融合的学者

开放词汇分割借助CLIP等大规模视觉-语言模型,突破了传统数据集预定义类别的限制,实现了跨场景的零样本理解。将此类能力扩展至3D分割面临挑战:CLIP基于图像的嵌入缺乏足够的几何细节。现有方法多引入额外分割模型或替换为微调过的变体,导致冗余或损失CLIP的通用语言能力。为此,我们提出SPNeRF,一种基于NeRF的零样本3D分割方法,利用几何先验。通过将3D场景中的几何基元融入NeRF训练,生成基元级的CLIP特征,避免点级特征的模糊性。同时,提出基于基元的合并机制,结合亲和度评分。无需额外分割模型,进一步挖掘了CLIP在3D分割中的潜力,在基准上实现显著优于原始LERF的效果。

原文摘要 · Abstract (English)

Open-vocabulary segmentation, powered by large visual-language models like CLIP, has expanded 2D segmentation capabilities beyond fixed classes predefined by the dataset, enabling zero-shot understanding across diverse scenes. Extending these capabilities to 3D segmentation introduces challenges, as CLIP's image-based embeddings often lack the geometric detail necessary for 3D scene segmentation. Recent methods tend to address this by introducing additional segmentation models or replacing CLIP with variations trained on segmentation data, which lead to redundancy or loss on CLIP's general language capabilities. To overcome this limitation, we introduce SPNeRF, a NeRF based zero-shot 3D segmentation approach that leverages geometric priors. We integrate geometric primitives derived from the 3D scene into NeRF training to produce primitive-wise CLIP features, avoiding the ambiguity of point-wise features. Additionally, we propose a primitive-based merging mechanism enhanced with affinity scores. Without relying on additional segmentation models, our method further explores CLIP's capability for 3D segmentation and achieves notable improvements over original LERF.

3D分割零样本视觉语言模型几何先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。