用语义引导的掩码策略提升点云自监督学习效果
Exploring Semantic Masked Autoencoder for Self-supervised Point Cloud Understanding
- 基于可学习原型建模物体组件语义,指导掩码策略
- 在ScanObjectNN等数据集上显著优于随机掩码方法
- 适合点云理解、3D目标识别等下游任务研究者
点云理解旨在从无标签数据中获取鲁棒且通用的特征表示。基于掩码点建模的方法在多个下游任务中表现出色。然而,现有预训练方法依赖随机掩码策略,通过恢复受损点云输入来建立感知能力,导致自监督模型难以捕捉合理的语义关系。为此,本文提出语义掩码自编码器(Semantic Masked Autoencoder),包含两个核心组件:基于原型的语义建模模块与增强型语义掩码策略。具体而言,在语义建模模块中,设计组件语义引导机制,利用一组可学习原型来捕捉物体不同组件的语义信息;基于这些原型,构建组件语义增强型掩码策略,克服随机掩码在完整覆盖组件结构上的不足。此外,引入组件语义增强型提示调优策略,进一步利用原型提升预训练模型在下游任务中的表现。在ScanObjectNN、ModelNet40和ShapeNetPart等数据集上的大量实验验证了所提模块的有效性。
原文摘要 · Abstract (English)
Point cloud understanding aims to acquire robust and general feature representations from unlabeled data. Masked point modeling-based methods have recently shown significant performance across various downstream tasks. These pre-training methods rely on random masking strategies to establish the perception of point clouds by restoring corrupted point cloud inputs, which leads to the failure of capturing reasonable semantic relationships by the self-supervised models. To address this issue, we propose Semantic Masked Autoencoder, which comprises two main components: a prototype-based component semantic modeling module and a component semantic-enhanced masking strategy. Specifically, in the component semantic modeling module, we design a component semantic guidance mechanism to direct a set of learnable prototypes in capturing the semantics of different components from objects. Leveraging these prototypes, we develop a component semantic-enhanced masking strategy that addresses the limitations of random masking in effectively covering complete component structures. Furthermore, we introduce a component semantic-enhanced prompt-tuning strategy, which further leverages these prototypes to improve the performance of pre-trained models in downstream tasks. Extensive experiments conducted on datasets such as ScanObjectNN, ModelNet40, and ShapeNetPart demonstrate the effectiveness of our proposed modules.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。