Tango3D实现2D图像与3D点云的精细对齐,兼顾全局检索与局部对应。
Tango3D: Towards Alignment for Global and Local 2D-3D Correspondence

- 用几何感知视觉主干和3D VAE编码图像与点云,统一到共享空间
- 在多个数据集上实现像素级对齐,同时保持优异的全局检索性能
- 适合需要细粒度3D对齐的下游任务,如3D生成与理解
现有3D基础模型通常将点云对齐到冻结的视觉-语言空间(如CLIP),通过压缩3D形状为全局向量实现强跨模态检索,但无法建立细粒度的像素-点对应。为此,我们提出Tango3D,一个统一密集对应与全局检索的基础模型。采用几何感知2D视觉主干和预训练3D VAE,将图像编码为2D补丁,点云编码为3D标记,并映射到单一共享空间,实现局部像素-点对齐与全局语义对齐。为稳定密集与全局目标的联合学习,引入三阶段渐进式训练策略。实验表明,该模型成功实现对象级别的像素-点对齐,同时保持竞争力的全局检索能力,这是现有3D基础模型不具备的联合能力。通过建立细粒度对齐特征空间,Tango3D为纯几何3D标记注入丰富语义,为广泛的密集3D下游任务铺平道路。
原文摘要 · Abstract (English)
Existing 3D foundation models typically align point clouds to frozen vision-language spaces like CLIP, which achieve strong cross-modal retrieval by compressing 3D shape into a global vector. However, this global-only alignment cannot establish fine-grained pixel-to-point correspondence. To solve this, we present Tango3D, a foundation model that unifies dense correspondence and global retrieval. We use a geometry-aware 2D visual backbone and a pretrained 3D VAE to encode images into 2D patches and point clouds into 3D tokens. These are mapped into a single shared space to achieve both local pixel-to-point alignment and global semantic alignment. To stabilize the joint learning of dense and global objectives, we introduce a three-stage progressive training strategy. Experiments show our model successfully achieves object-level pixel-to-point alignment while maintaining competitive global retrieval, a joint capability not offered by existing 3D foundation models. By establishing a fine-grained alignment feature space, Tango3D injects rich semantics into purely geometric 3D tokens, paving the way for a wide range of dense 3D downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。