arXiv:2410.09519cs.CVcs.AI2024-10中稿 · ACML 2024被引 1

用图像与点云的对应关系实现3D数据自监督预训练

Pic@Point: Cross-Modal Learning by Local and Global Point-Picture Correspondence

  • 通过2D图像与3D点云的结构化对应构建对比学习信号
  • 在多个3D基准上超越现有自监督方法性能
  • 适合做3D视觉表征学习的研究者参考

自监督预训练在自然语言处理和2D视觉中取得显著成功,但在3D数据领域仍面临挑战。基于掩码重建的方法在非结构化点云上存在固有困难,而许多对比学习任务则缺乏复杂性和信息量。本文提出Pic@Point,一种基于2D-3D结构对应关系的高效对比学习方法。利用富含语义与上下文知识的图像线索,为不同抽象层级的点云表示提供引导信号。该轻量级方法在多个3D基准测试中优于当前最先进的自监督预训练方法。

原文摘要 · Abstract (English)

Self-supervised pre-training has achieved remarkable success in NLP and 2D vision. However, these advances have yet to translate to 3D data. Techniques like masked reconstruction face inherent challenges on unstructured point clouds, while many contrastive learning tasks lack in complexity and informative value. In this paper, we present Pic@Point, an effective contrastive learning method based on structural 2D-3D correspondences. We leverage image cues rich in semantic and contextual knowledge to provide a guiding signal for point cloud representations at various abstraction levels. Our lightweight approach outperforms state-of-the-art pre-training methods on several 3D benchmarks.

3D视觉对比学习自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。