arXiv:2409.13982cs.CVcs.MM2024-09

用噪声过滤提升3D分割精度,无需标注即可识别物体。

CUS3D :CLIP-based Unsupervised 3D Segmentation via Object-level Denoise

  • 在2D到3D特征映射中加入物体级去噪模块,清除噪声干扰。
  • 在无监督和开放词汇分割任务上均超越现有方法,效果显著。
  • 适合对零样本3D语义分割感兴趣的开发者和研究者。

为缓解3D数据标注成本高的问题,现有方法常采用无监督、开放词汇的语义分割,借助2D CLIP的语义知识。本文针对以往研究忽略从2D到3D特征投影过程中引入的“噪声”问题,提出一种新型蒸馏学习框架CUS3D。其核心包括:一个物体级去噪投影模块,用于筛选并去除噪声,确保更准确的3D特征;一个基于多模态的蒸馏学习模块,在以物体为中心的约束下,将3D特征与CLIP语义空间对齐,实现先进的无监督语义分割。我们在无监督和开放词汇分割任务上进行了全面实验,结果一致表明,该模型在无监督分割上表现优异,并在开放词汇分割中展现出强有效性。

原文摘要 · Abstract (English)

To ease the difficulty of acquiring annotation labels in 3D data, a common method is using unsupervised and open-vocabulary semantic segmentation, which leverage 2D CLIP semantic knowledge. In this paper, unlike previous research that ignores the ``noise'' raised during feature projection from 2D to 3D, we propose a novel distillation learning framework named CUS3D. In our approach, an object-level denosing projection module is designed to screen out the ``noise'' and ensure more accurate 3D feature. Based on the obtained features, a multimodal distillation learning module is designed to align the 3D feature with CLIP semantic feature space with object-centered constrains to achieve advanced unsupervised semantic segmentation. We conduct comprehensive experiments in both unsupervised and open-vocabulary segmentation, and the results consistently showcase the superiority of our model in achieving advanced unsupervised segmentation results and its effectiveness in open-vocabulary segmentation.

3D分割无监督学习CLIP去噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。