arXiv:2506.21724cs.CV2025-06NeurIPS

通过隐空间预测统一掩码建模与不变性学习,提升3D点云自监督表征能力。

Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning

  • 在隐空间进行预测,避免输入空间重建限制语义捕捉。
  • 在ScanObjectNN上达90.53%准确率,预训练93万形状后达93.72%。
  • 设计不对称结构与多掩码采样,防止形状信息泄露。

从无结构的3D点云中学习具有语义意义的表征仍是计算机视觉中的核心挑战,尤其在缺乏大规模标注数据集的情况下。尽管掩码点建模(MPM)广泛应用于自监督3D学习,但其基于重建的目标可能限制对高层语义的捕捉。我们提出AsymDSD,一种不对称双自蒸馏框架,通过在隐空间而非输入空间进行预测,统一掩码建模与不变性学习。AsymDSD基于联合嵌入架构,引入多项关键设计:高效不对称设置、禁用掩码查询间的注意力以防止形状泄露、多掩码采样以及点云版多裁剪。该方法在ScanObjectNN上取得90.53%的当前最优结果,当在93万形状上预训练时进一步提升至93.72%,超越已有方法。

原文摘要 · Abstract (English)

Learning semantically meaningful representations from unstructured 3D point clouds remains a central challenge in computer vision, especially in the absence of large-scale labeled datasets. While masked point modeling (MPM) is widely used in self-supervised 3D learning, its reconstruction-based objective can limit its ability to capture high-level semantics. We propose AsymDSD, an Asymmetric Dual Self-Distillation framework that unifies masked modeling and invariance learning through prediction in the latent space rather than the input space. AsymDSD builds on a joint embedding architecture and introduces several key design choices: an efficient asymmetric setup, disabling attention between masked queries to prevent shape leakage, multi-mask sampling, and a point cloud adaptation of multi-crop. AsymDSD achieves state-of-the-art results on ScanObjectNN (90.53%) and further improves to 93.72% when pretrained on 930k shapes, surpassing prior methods.

3D表征自监督点云蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。