arXiv:2503.05283cs.CV2025-03CVPR被引 6

探索3D与文本特征空间的对齐,发现降维子空间能显著提升匹配准确率

Escaping Plato's Cave: Towards the Alignment of 3D and Text Latent Spaces

  • 通过投影到低维子空间实现3D与文本特征对齐
  • 在图像检索任务中准确率提升显著,优于直接对齐方法
  • 适合研究多模态对齐、3D表示学习的学者参考

近期研究表明,大规模训练下单模态2D视觉与文本编码器虽来自不同表征,却涌现出相似的结构特性。然而,3D编码器与其他模态的关系仍不明确。现有3D基础模型通常依赖固定编码器进行显式对齐训练。本文研究了单模态3D编码器与文本特征空间的后训练对齐可能性。结果表明,简单的后训练特征对齐性能有限。我们转而分析特征空间的子结构,发现将学习到的表征投影至精心选择的低维子空间后,对齐质量显著提升,在匹配与检索任务中表现更优。进一步分析揭示这些共享子空间大致分离了语义与几何信息。本工作首次建立3D与文本特征空间后训练对齐的基线,揭示了3D数据与其它表征的共性与差异。代码与权重已开源。

原文摘要 · Abstract (English)

Recent works have shown that, when trained at scale, uni-modal 2D vision and text encoders converge to learned features that share remarkable structural properties, despite arising from different representations. However, the role of 3D encoders with respect to other modalities remains unexplored. Furthermore, existing 3D foundation models that leverage large datasets are typically trained with explicit alignment objectives with respect to frozen encoders from other representations. In this work, we investigate the possibility of a posteriori alignment of representations obtained from uni-modal 3D encoders compared to text-based feature spaces. We show that naive post-training feature alignment of uni-modal text and 3D encoders results in limited performance. We then focus on extracting subspaces of the corresponding feature spaces and discover that by projecting learned representations onto well-chosen lower-dimensional subspaces the quality of alignment becomes significantly higher, leading to improved accuracy on matching and retrieval tasks. Our analysis further sheds light on the nature of these shared subspaces, which roughly separate between semantic and geometric data representations. Overall, ours is the first work that helps to establish a baseline for post-training alignment of 3D uni-modal and text feature spaces, and helps to highlight both the shared and unique properties of 3D data compared to other representations. Our code and weights are available at https://github.com/Souhail-01/3d-text-alignment

3D对齐特征空间多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。