提出CMCR框架,让3D表示同时学好跨模态共性和特有特征。
Is Contrastive Distillation Enough for Learning Comprehensive 3D Representations?
- 引入掩码图像建模与占据估计,增强模态特有特征学习
- 设计多模态统一代码本,实现跨模态嵌入空间共享
- 在下游任务中优于现有图像到激光雷达对比蒸馏方法
跨模态对比蒸馏近年被用于学习有效的3D表示,但现有方法主要关注模态共享特征,忽视了模态特有特征的预训练,导致表示性能受限。本文理论分析了现有对比方法在3D表示学习中的局限性,提出新框架CMCR(跨模态综合表示学习)以解决此问题。该方法通过引入掩码图像建模和占据估计任务,引导网络学习更全面的模态特有特征;提出新颖的多模态统一代码本,学习跨模态共享的嵌入空间;并设计几何增强的掩码图像建模进一步提升3D表示能力。大量实验表明,该方法有效缓解传统方法的挑战,在下游任务中持续优于现有图像到激光雷达对比蒸馏方法。代码将发布于https://github.com/Eaphan/CMCR。
原文摘要 · Abstract (English)
Cross-modal contrastive distillation has recently been explored for learning effective 3D representations. However, existing methods focus primarily on modality-shared features, neglecting the modality-specific features during the pre-training process, which leads to suboptimal representations. In this paper, we theoretically analyze the limitations of current contrastive methods for 3D representation learning and propose a new framework, namely CMCR (Cross-Modal Comprehensive Representation Learning), to address these shortcomings. Our approach improves upon traditional methods by better integrating both modality-shared and modality-specific features. Specifically, we introduce masked image modeling and occupancy estimation tasks to guide the network in learning more comprehensive modality-specific features. Furthermore, we propose a novel multi-modal unified codebook that learns an embedding space shared across different modalities. Besides, we introduce geometry-enhanced masked image modeling to further boost 3D representation learning. Extensive experiments demonstrate that our method mitigates the challenges faced by traditional approaches and consistently outperforms existing image-to-LiDAR contrastive distillation methods in downstream tasks. Code will be available at https://github.com/Eaphan/CMCR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。