arXiv:2509.01250cs.CV2025-09ICCV被引 10

通过双视角交叉重建提升点云自监督学习的多样性与挑战性

Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views

  • 设计双视角解耦生成与交叉重建机制
  • 在ScanObjectNN上比基线提升6.5%~7.0%
  • 适合追求更高性能的3D点云自监督研究者

点云自监督学习因其在众多应用中的潜力而受到广泛关注。现有生成式方法多聚焦于单视角内被遮掩点的恢复。本文提出双视角预训练范式,显著增加数据多样性和差异性,从而带来更具挑战性的预训练任务。我们提出Point-PQAE,先生成两个解耦的点云视角,再相互重建。首次引入点云视图裁剪机制,并设计新型位置编码以表征两视角间的3D相对位置。相比自重建,交叉重建大幅提高预训练难度,使本方法在三个ScanObjectNN变体上分别超越基线(Point-MAE)6.5%、7.0%和6.7%(采用Mlp-Linear评估协议)。代码已开源。

原文摘要 · Abstract (English)

Point cloud learning, especially in a self-supervised way without manual labels, has gained growing attention in both vision and learning communities due to its potential utility in a wide range of applications. Most existing generative approaches for point cloud self-supervised learning focus on recovering masked points from visible ones within a single view. Recognizing that a two-view pre-training paradigm inherently introduces greater diversity and variance, it may thus enable more challenging and informative pre-training. Inspired by this, we explore the potential of two-view learning in this domain. In this paper, we propose Point-PQAE, a cross-reconstruction generative paradigm that first generates two decoupled point clouds/views and then reconstructs one from the other. To achieve this goal, we develop a crop mechanism for point cloud view generation for the first time and further propose a novel positional encoding to represent the 3D relative position between the two decoupled views. The cross-reconstruction significantly increases the difficulty of pre-training compared to self-reconstruction, which enables our method to surpass previous single-modal self-reconstruction methods in 3D self-supervised learning. Specifically, it outperforms the self-reconstruction baseline (Point-MAE) by 6.5%, 7.0%, and 6.7% in three variants of ScanObjectNN with the Mlp-Linear evaluation protocol. The code is available at https://github.com/aHapBean/Point-PQAE.

点云学习自监督交叉重建三维视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。