arXiv:2603.25165cs.CV2026-03中稿 · CVPR被引 3

提出点云自监督学习新框架,提升3D场景实例感知能力。

Towards Foundation Models for 3D Scene Understanding: Instance-Aware Self-Supervised Learning for Point Clouds

  • 通过几何感知分支联合学习语义与空间关系
  • 在五个数据集上实例分割平均提升3.5% mAP
  • 适合构建可迁移的3D基础模型,尤其关注实例定位

点云自监督学习近年显著提升了无需人工标注的3D场景理解能力。现有方法侧重语义一致性或掩码场景建模,但所得表征在实例定位任务上迁移性差,通常需全量微调才能达优性能。实例感知是3D感知的核心,弥合此差距对实现支持各类下游任务的3D基础模型至关重要。本文提出PointINS,一种面向实例的自监督框架,通过几何感知学习增强点云表征。该框架引入正交偏移分支,联合学习高层语义与几何推理,实现实例感知。我们识别出两个对鲁棒实例定位至关重要的性质,并构建互补正则化策略:偏移分布正则化(ODR)使预测偏移匹配经验几何先验;空间聚类正则化(SCR)通过伪实例掩码约束偏移,强化局部一致性。在五个数据集上的大量实验表明,PointINS在室内实例分割上平均提升3.5% mAP,室外全景分割提升4.1% PQ,为可扩展3D基础模型铺平道路。

原文摘要 · Abstract (English)

Recent advances in self-supervised learning (SSL) for point clouds have substantially improved 3D scene understanding without human annotations. Existing approaches emphasize semantic awareness by enforcing feature consistency across augmented views or by masked scene modeling. However, the resulting representations transfer poorly to instance localization, and often require full finetuning for strong performance. Instance awareness is a fundamental component of 3D perception, thus bridging this gap is crucial for progressing toward true 3D foundation models that support all downstream tasks on 3D data. In this work, we introduce PointINS, an instance-oriented self-supervised framework that enriches point cloud representations through geometry-aware learning. PointINS employs an orthogonal offset branch to jointly learn high-level semantic understanding and geometric reasoning, yielding instance awareness. We identify two consistent properties essential for robust instance localization and formulate them as complementary regularization strategies, Offset Distribution Regularization (ODR), which aligns predicted offsets with empirically observed geometric priors, and Spatial Clustering Regularization (SCR), which enforces local coherence by regularizing offsets with pseudo-instance masks. Through extensive experiments across five datasets, PointINS achieves on average +3.5% mAP improvement for indoor instance segmentation and +4.1% PQ gain for outdoor panoptic segmentation, paving the way for scalable 3D foundation models.

3D感知自监督学习点云实例分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。