通过跨域专家协作,提升3D点云自监督学习的表征能力。
Point Cloud Mixture-of-Domain-Experts Model for 3D Self-supervised Learning
- 构建场景与物体域专家混合模型,融合跨域知识。
- 设计块到场景预训练策略,同时优化物体与场景表征。
- 在下游任务中表现优于现有方法,适合3D视觉研究者。
点云作为三维数据的主要表示形式,可分为场景域点云和物体域点云。点云自监督学习(SSL)已成为学习三维表征的主流范式。然而,现有方法主要关注单一域内的领域特定表征学习,忽视了跨域知识的互补性,限制了三维表征的学习效果。本文提出一种基于块到场景预训练策略的点云域专家混合模型(Point-MoDE)。首先,构建由场景域专家和多个共享物体域专家组成的混合专家模型;其次,提出块到场景预训练策略,利用物体域中的点块特征,通过物体级块掩码重建和场景级块位置回归,预测其在场景域中的初始位置。该策略通过整合物体与场景间的互补知识,同时促进物体域与场景域表征的学习,获得更全面的三维表征。大量下游任务实验表明,本模型具有显著优势。
原文摘要 · Abstract (English)
Point clouds, as a primary representation of 3D data, can be categorized into scene domain point clouds and object domain point clouds. Point cloud self-supervised learning (SSL) has become a mainstream paradigm for learning 3D representations. However, existing point cloud SSL primarily focuses on learning domain-specific 3D representations within a single domain, neglecting the complementary nature of cross-domain knowledge, which limits the learning of 3D representations. In this paper, we propose to learn a comprehensive Point cloud Mixture-of-Domain-Experts model (Point-MoDE) via a block-to-scene pre-training strategy. Specifically, we first propose a mixture-of-domain-expert model consisting of scene domain experts and multiple shared object domain experts. Furthermore, we propose a block-to-scene pretraining strategy, which leverages the features of point blocks in the object domain to regress their initial positions in the scene domain through object-level block mask reconstruction and scene-level block position regression. By integrating the complementary knowledge between object and scene, this strategy simultaneously facilitates the learning of both object-domain and scene-domain representations, leading to a more comprehensive 3D representation. Extensive experiments in downstream tasks demonstrate the superiority of our model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。