arXiv:2601.19247cs.CV2026-01

通过解耦3D高斯点云特征,实现文本、图像与3D场景的精准对齐。

TIGaussian: Disentangle Gaussians for Spatial-Awared Text-Image-3D Alignment

论文配图:TIGaussian: Disentangle Gaussians for Spatial-Awared Text-Image-3D Alignment
图 1 · 摘自论文原文
  • 将3D高斯结构拆解为紧凑潜在表示,提升特征泛化能力。
  • 多视角融合+扩散先验,解决图像-3D对齐中的视角歧义问题。
  • 适合做跨模态3D理解、零样本分类和场景识别的研究者使用。

尽管视觉语言模型已显著拉近文本与图像间的特征关联,但引入点云与3D高斯等3D模态数据,进一步推动了3D相关任务的预训练发展,如跨模态检索、零样本分类和场景识别。针对3D特征提取困难及模态间鸿沟问题,本文提出TIGaussian框架,利用3D高斯喷溅(3DGS)特性,通过多分支3DGS分词器与模态特定的3D特征对齐策略增强跨模态对齐。具体而言,多分支3DGS分词器将3DGS结构的内在属性解耦为紧凑潜在表示,实现更通用的特征提取;为弥合模态差距,设计双向跨模态对齐机制:多视角特征融合利用扩散先验缓解图像-3D对齐中的视角模糊性,而文本-3D投影模块则自适应地将3D特征映射至文本嵌入空间,提升文本-3D对齐效果。在多个数据集上的大量实验表明,TIGaussian在多项任务中达到当前最优性能。

原文摘要 · Abstract (English)

While visual-language models have profoundly linked features between texts and images, the incorporation of 3D modality data, such as point clouds and 3D Gaussians, further enables pretraining for 3D-related tasks, e.g., cross-modal retrieval, zero-shot classification, and scene recognition. As challenges remain in extracting 3D modal features and bridging the gap between different modalities, we propose TIGaussian, a framework that harnesses 3D Gaussian Splatting (3DGS) characteristics to strengthen cross-modality alignment through multi-branch 3DGS tokenizer and modality-specific 3D feature alignment strategies. Specifically, our multi-branch 3DGS tokenizer decouples the intrinsic properties of 3DGS structures into compact latent representations, enabling more generalizable feature extraction. To further bridge the modality gap, we develop a bidirectional cross-modal alignment strategies: a multi-view feature fusion mechanism that leverages diffusion priors to resolve perspective ambiguity in image-3D alignment, while a text-3D projection module adaptively maps 3D features to text embedding space for better text-3D alignment. Extensive experiments on various datasets demonstrate the state-of-the-art performance of TIGaussian in multiple tasks.

3D生成跨模态对齐高斯溅射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。