用双曲空间建模3D点云的层次结构,提升跨模态迁移能力。
Hyperbolic Contrastive Learning for Hierarchical 3D Point Cloud Embedding
- 在双曲空间中设计多模态对比学习框架,统一文本、图像与3D点云
- 引入蕴含、模态差距与对齐正则项,学习跨模态层次关系
- 显著提升3D点云下游任务性能,适合多模态场景下的表征学习
双曲空间能更高效地建模复杂层次结构,对多模态数据尤其有益。尽管双曲几何在语言-图像预训练中表现优异,其在融合语言、图像与3D点云三者方面的潜力仍待探索。本文将3D点云模态拓展至双曲多模态对比预训练框架中,并引入蕴含、模态差距和对齐正则项,以学习各模态内部及跨文本、2D图像与3D点云之间的层次结构。实验表明,所提训练策略生成的3D点云编码器性能优异,获得的层次化嵌入显著提升多个下游任务表现。
原文摘要 · Abstract (English)
Hyperbolic spaces allow for more efficient modeling of complex, hierarchical structures, which is particularly beneficial in tasks involving multi-modal data. Although hyperbolic geometries have been proven effective for language-image pre-training, their capabilities to unify language, image, and 3D Point Cloud modalities are under-explored. We extend the 3D Point Cloud modality in hyperbolic multi-modal contrastive pre-training. Additionally, we explore the entailment, modality gap, and alignment regularizers for learning hierarchical 3D embeddings and facilitating the transfer of knowledge from both Text and Image modalities. These regularizers enable the learning of intra-modal hierarchy within each modality and inter-modal hierarchy across text, 2D images, and 3D Point Clouds. Experimental results demonstrate that our proposed training strategy yields an outstanding 3D Point Cloud encoder, and the obtained 3D Point Cloud hierarchical embeddings significantly improve performance on various downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。