多模态多任务预训练提升点云理解能力
Multi-modal Multi-task Pre-training for Improved Point Cloud Understanding
- 设计三种自监督任务联合优化点云与图像特征
- 在多个下游任务上超越现有方法,无需3D标注
- 适合需要强泛化能力的3D视觉研究者
近期多模态预训练方法通过对齐3D形状与其对应2D图像的多模态特征,在学习3D表示方面展现出良好效果。然而,现有框架主要依赖单一预训练任务来获取多模态数据,限制了模型从其他相关任务中汲取丰富信息的能力,从而影响其在复杂多样的下游任务中的表现。为此,我们提出MMPT框架,一种用于增强点云理解的多模态多任务预训练方法。具体设计三个预训练任务:(i) Token-level reconstruction (TLR) 旨在恢复被掩码的点令牌,赋予模型表征学习能力;(ii) Point-level reconstruction (PLR) 直接预测被掩码点的位置,重建后的点云可作为后续任务的变换输入;(iii) Multi-modal contrastive learning (MCL) 联合模态内与跨模态的特征对应关系,以自监督方式整合3D点云与2D图像模态的丰富学习信号。该框架无需任何3D标注,具备在大规模数据集上扩展的能力。训练后的编码器可有效迁移到多种下游任务。为验证有效性,我们在多个判别性与生成性应用中,基于广泛使用的基准测试,将该方法与现有最先进方法进行了对比。
原文摘要 · Abstract (English)
Recent advances in multi-modal pre-training methods have shown promising effectiveness in learning 3D representations by aligning multi-modal features between 3D shapes and their corresponding 2D counterparts. However, existing multi-modal pre-training frameworks primarily rely on a single pre-training task to gather multi-modal data in 3D applications. This limitation prevents the models from obtaining the abundant information provided by other relevant tasks, which can hinder their performance in downstream tasks, particularly in complex and diverse domains. In order to tackle this issue, we propose MMPT, a Multi-modal Multi-task Pre-training framework designed to enhance point cloud understanding. Specifically, three pre-training tasks are devised: (i) Token-level reconstruction (TLR) aims to recover masked point tokens, endowing the model with representative learning abilities. (ii) Point-level reconstruction (PLR) is integrated to predict the masked point positions directly, and the reconstructed point cloud can be considered as a transformed point cloud used in the subsequent task. (iii) Multi-modal contrastive learning (MCL) combines feature correspondences within and across modalities, thus assembling a rich learning signal from both 3D point cloud and 2D image modalities in a self-supervised manner. Moreover, this framework operates without requiring any 3D annotations, making it scalable for use with large datasets. The trained encoder can be effectively transferred to various downstream tasks. To demonstrate its effectiveness, we evaluated its performance compared to state-of-the-art methods in various discriminant and generative applications under widely-used benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。