用图像、点云和文本三模态重建,提升3D形状检索精度
Enhanced Cross-modal 3D Retrieval via Tri-modal Reconstruction
- 融合多视图图像、点云与文本,实现三模态对齐
- 在Text2Shape数据集上显著超越现有方法,双方向检索均更优
- 适合做3D内容检索、跨模态理解的研究者参考
跨模态3D检索是一项关键但具有挑战性的任务,旨在实现3D与文本之间的双向检索。当前方法主要依赖某种3D表示(如点云),很少利用2D-3D的一致性与互补关系,限制了性能。为此,我们提出联合使用多视图图像与点云来表示3D形状,促进图像、点、文本三模态对齐,以增强跨模态3D检索。特别地,引入三模态重建机制以提升编码器的泛化能力:给定点特征,在文本特征引导下重建图像特征,反之亦然。通过精细的2D-3D融合,将对齐后的点云与多视图图像特征聚合为多模态嵌入,增强几何与语义理解。针对现有数据集中大量3D形状与文本语义相似的问题,采用难负样本对比训练,强化对更具区分度的难负样本的学习,生成鲁棒的判别性嵌入。在Text2Shape数据集上的大量实验表明,该方法在形状到文本与文本到形状检索任务中均显著优于现有最先进方法。
原文摘要 · Abstract (English)
Cross-modal 3D retrieval is a critical yet challenging task, aiming to achieve bi-directional retrieval between 3D and text modalities. Current methods predominantly rely on a certain 3D representation (e.g., point cloud), with few exploiting the 2D-3D consistency and complementary relationships, which constrains their performance. To bridge this gap, we propose to adopt multi-view images and point clouds to jointly represent 3D shapes, facilitating tri-modal alignment (i.e., image, point, text) for enhanced cross-modal 3D retrieval. Notably, we introduce tri-modal reconstruction to improve the generalization ability of encoders. Given point features, we reconstruct image features under the guidance of text features, and vice versa. With well-aligned point cloud and multi-view image features, we aggregate them as multimodal embeddings through fine-grained 2D-3D fusion to enhance geometric and semantic understanding. Recognizing the significant noise in current datasets where many 3D shapes and texts share similar semantics, we employ hard negative contrastive training to emphasize harder negatives with greater significance, leading to robust discriminative embeddings. Extensive experiments on the Text2Shape dataset demonstrate that our method significantly outperforms previous state-of-the-art methods in both shape-to-text and text-to-shape retrieval tasks by a substantial margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。