用3D高斯点云统一建模图文与3D,提升多模态理解能力。
UniGS: Unified Language-Image-3D Pretraining with Gaussian Splatting
- 用3D高斯点云替代点云,实现更精细的3D场景表示。
- 在多个数据集上零样本分类提升9.36%,文本检索提升4.3%。
- 适合做多模态3D理解、跨模态对齐的研究者和开发者。
近期多模态3D预训练方法在学习文本、图像与点云联合表征方面展现出良好效果。然而,采用点云作为3D表示难以充分捕捉3D世界的复杂性,且离散点与密集2D像素间存在明显鸿沟。为此,我们提出UniGS,将3D高斯点云(3DGS)引入多模态预训练以增强3D表征能力。首先,利用3DGS将3D场景建模为带颜色与不透明度的3D高斯分布,融合完整场景信息并建立与2D图像的强关联。接着,基于预训练视觉-语言模型,通过真实世界图像-文本对构建共享的视觉与文本空间。随后,采用3D编码器将优化后的3DGS与语言-图像表征对齐,学习统一的多模态表征。为促进3D编码器提取全局显式3D特征并实现更好跨模态对齐,我们引入新颖的高斯感知引导模块,指导3D领域细粒度表征学习。在Objaverse、ABO、MVImgNet和SUN RGBD数据集上的广泛实验表明,该方法在零样本分类、文本驱动检索与开放世界理解任务中均取得显著效果,性能超越当前最优方法Uni3D,其中零样本分类提升9.36%,文本检索提升4.3%,开放世界理解提升7.92%。
原文摘要 · Abstract (English)
Recent advancements in multi-modal 3D pre-training methods have shown promising efficacy in learning joint representations of text, images, and point clouds. However, adopting point clouds as 3D representation fails to fully capture the intricacies of the 3D world and exhibits a noticeable gap between the discrete points and the dense 2D pixels of images. To tackle this issue, we propose UniGS, integrating 3D Gaussian Splatting (3DGS) into multi-modal pre-training to enhance the 3D representation. We first rely on the 3DGS representation to model the 3D world as a collection of 3D Gaussians with color and opacity, incorporating all the information of the 3D scene while establishing a strong connection with 2D images. Then, to achieve Language-Image-3D pertaining, UniGS starts with a pre-trained vision-language model to establish a shared visual and textual space through extensive real-world image-text pairs. Subsequently, UniGS employs a 3D encoder to align the optimized 3DGS with the Language-Image representations to learn unified multi-modal representations. To facilitate the extraction of global explicit 3D features by the 3D encoder and achieve better cross-modal alignment, we additionally introduce a novel Gaussian-Aware Guidance module that guides the learning of fine-grained representations of the 3D domain. Through extensive experiments across the Objaverse, ABO, MVImgNet and SUN RGBD datasets with zero-shot classification, text-driven retrieval and open-world understanding tasks, we demonstrate the effectiveness of UniGS in learning a more general and stronger aligned multi-modal representation. Specifically, UniGS achieves leading results across different 3D tasks with remarkable improvements over previous SOTA, Uni3D, including on zero-shot classification (+9.36%), text-driven retrieval (+4.3%) and open-world understanding (+7.92%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。