arXiv:2409.15803cs.CV2024-09被引 12

3D-JEPA通过上下文感知解码器提升3D自监督表征学习效果

3D-JEPA: A Joint Embedding Predictive Architecture for 3D Self-Supervised Representation Learning

  • 采用多块采样策略生成上下文与目标块,增强信息表达
  • 上下文感知解码器使模型聚焦语义建模而非记忆细节,准确率达88.65%
  • 无需生成重建,适合高效预训练,适用于各类下游任务

基于不变性与生成式的方法在3D自监督表征学习(SSRL)中表现突出。然而,前者依赖人工设计的数据增强,引入偏差且不适用于所有下游任务;后者无差别重建掩码区域,导致无关细节进入表示空间。为此,我们提出3D-JEPA,一种新型非生成式3D SSRL框架。具体而言,设计多块采样策略,生成充分信息的上下文块和多个代表性目标块;提出上下文感知解码器,持续输入上下文信息,促使编码器学习语义建模而非记忆目标块相关上下文。整体上,3D-JEPA通过编码器与上下文感知解码器架构,从上下文块预测目标块表示。在不同数据集上的多种下游任务验证了其有效性与高效性,仅用150个预训练轮次即达到PB_T50_RS上88.65%的准确率。

原文摘要 · Abstract (English)

Invariance-based and generative methods have shown a conspicuous performance for 3D self-supervised representation learning (SSRL). However, the former relies on hand-crafted data augmentations that introduce bias not universally applicable to all downstream tasks, and the latter indiscriminately reconstructs masked regions, resulting in irrelevant details being saved in the representation space. To solve the problem above, we introduce 3D-JEPA, a novel non-generative 3D SSRL framework. Specifically, we propose a multi-block sampling strategy that produces a sufficiently informative context block and several representative target blocks. We present the context-aware decoder to enhance the reconstruction of the target blocks. Concretely, the context information is fed to the decoder continuously, facilitating the encoder in learning semantic modeling rather than memorizing the context information related to target blocks. Overall, 3D-JEPA predicts the representation of target blocks from a context block using the encoder and context-aware decoder architecture. Various downstream tasks on different datasets demonstrate 3D-JEPA's effectiveness and efficiency, achieving higher accuracy with fewer pretraining epochs, e.g., 88.65% accuracy on PB_T50_RS with 150 pretraining epochs.

3D表征学习自监督上下文感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。