arXiv:2512.11465cs.CVcs.LG2025-12AAAI被引 2

通过可观察点的软图蒸馏,提升3D点云自监督表征能力。

DOS: Distilling Observable Softmaps of Zipfian Prototypes for Self-Supervised Point Representation

论文配图:DOS: Distilling Observable Softmaps of Zipfian Prototypes for Self-Supervised Point Representation
图 1 · 摘自论文原文
  • 仅在可见点上蒸馏语义软图,避免掩码区域信息泄露。
  • 在多个基准上超越当前最优方法,无需额外数据或标注。
  • 引入幂律原型与改进的Sinkhorn算法,解决语义分布不均问题。

近期自监督学习(SSL)在无需人工标注的情况下,展现出学习3D点云表征的巨大潜力。然而,由于几何不规则、重建捷径以及语义分布不均等问题,3D点云的自监督学习仍面临严峻挑战。本文提出DOS(Distilling Observable Softmaps),一种新型自监督框架,仅在可观察(未被掩码)点上自蒸馏语义相关性软图。该策略有效防止了掩码区域的信息泄漏,并提供比离散令牌-原型分配更丰富的监督信号。为应对无监督场景下的语义分布不均问题,我们引入幂律原型(Zipfian prototypes),并采用改进的Sinkhorn-Knopp算法——Zipf-Sinkhorn,对原型使用施加幂律先验,并在训练过程中调节目标软图的锐度。DOS在nuScenes、Waymo、SemanticKITTI、ScanNet和ScanNet200等多个基准上,在语义分割和3D物体检测任务中均优于现有最先进方法,且无需依赖额外数据或标注。结果表明,可观察点软图蒸馏为学习鲁棒3D表示提供了一种可扩展且高效的新范式。

原文摘要 · Abstract (English)

Recent advances in self-supervised learning (SSL) have shown tremendous potential for learning 3D point cloud representations without human annotations. However, SSL for 3D point clouds still faces critical challenges due to irregular geometry, shortcut-prone reconstruction, and unbalanced semantics distribution. In this work, we propose DOS (Distilling Observable Softmaps), a novel SSL framework that self-distills semantic relevance softmaps only at observable (unmasked) points. This strategy prevents information leakage from masked regions and provides richer supervision than discrete token-to-prototype assignments. To address the challenge of unbalanced semantics in an unsupervised setting, we introduce Zipfian prototypes and incorporate them using a modified Sinkhorn-Knopp algorithm, Zipf-Sinkhorn, which enforces a power-law prior over prototype usage and modulates the sharpness of the target softmap during training. DOS outperforms current state-of-the-art methods on semantic segmentation and 3D object detection across multiple benchmarks, including nuScenes, Waymo, SemanticKITTI, ScanNet, and ScanNet200, without relying on extra data or annotations. Our results demonstrate that observable-point softmaps distillation offers a scalable and effective paradigm for learning robust 3D representations.

3D点云自监督学习语义分割原型蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。