arXiv:2409.15810cs.CV2024-09中稿 · IROS2024被引 5

用双曲空间提升3D点云与图像的对比学习,分类性能显著提高。

Hyperbolic Image-and-Pointcloud Contrastive Learning for 3D Classification

  • 在双曲空间中构建图像与点云的对比学习框架,捕捉层级语义关系。
  • 在ScanObjectNN上提升物体分类准确率2.8%,少样本分类提升5.9%。
  • 适合做3D多模态表示学习、点云分类及需要层次结构建模的任务。

3D对比表示学习在各类下游任务中表现出色。然而,现有基于余弦相似度的对比学习方法在欧氏空间中难以深入挖掘多模态数据内部的层级结构及跨模态语义关联。为此,本文提出双曲图像-点云对比学习方法(HyperIPC)。在模态内分支中,利用点云的内在几何结构,在双曲空间中学习其嵌入表示,以捕捉不变特征;在跨模态分支中,借助图像引导点云建立强语义层级关联。实验表明,HyperIPC表现优异:在ScanObjectNN数据集上,物体分类准确率相比基线提升2.8%,少样本分类性能提升5.9%。消融实验与验证测试进一步证实了模型参数设置的合理性及各模块的有效性。

原文摘要 · Abstract (English)

3D contrastive representation learning has exhibited remarkable efficacy across various downstream tasks. However, existing contrastive learning paradigms based on cosine similarity fail to deeply explore the potential intra-modal hierarchical and cross-modal semantic correlations about multi-modal data in Euclidean space. In response, we seek solutions in hyperbolic space and propose a hyperbolic image-and-pointcloud contrastive learning method (HyperIPC). For the intra-modal branch, we rely on the intrinsic geometric structure to explore the hyperbolic embedding representation of point cloud to capture invariant features. For the cross-modal branch, we leverage images to guide the point cloud in establishing strong semantic hierarchical correlations. Empirical experiments underscore the outstanding classification performance of HyperIPC. Notably, HyperIPC enhances object classification results by 2.8% and few-shot classification outcomes by 5.9% on ScanObjectNN compared to the baseline. Furthermore, ablation studies and confirmatory testing validate the rationality of HyperIPC's parameter settings and the effectiveness of its submodules.

3D分类对比学习双曲空间多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。