arXiv:2607.05906cs.CV2026-07被引 1

用图文联合监督提升3D高斯表示的语义理解能力。

GaussFusion: Towards Multimodal 3D Gaussian Pretraining

论文配图:GaussFusion: Towards Multimodal 3D Gaussian Pretraining
图 1 · 摘自论文原文
  • 通过跨模态对齐,融合图像与文本信息增强高斯编码器
  • 多尺度空洞掩码机制使模型同时捕捉局部细节与整体结构
  • 在多个下游任务中表现优于现有方法,适合3D生成与理解场景

3D高斯点云渲染提供了一种显式建模几何与外观的表示方式,是3D表征学习的可扩展基础。现有高斯表示的预训练方法(如掩码高斯重建)主要捕捉局部结构,缺乏语义监督。本文提出GaussFusion,一种面向3D高斯表示的多模态预训练框架。该框架通过跨模态语义对齐,将图像与文本监督引入掩码高斯建模,使高斯编码器在预训练阶段同时学习视觉与语言级语义信息。为适应高斯原始点分布不均的特点,进一步提出基于高斯显著性的多尺度空洞掩码(GSHM),根据高斯显著性构建空间连续的掩码区域,并在多尺度上应用空洞掩码,促使编码器捕获细粒度局部模式与更广泛的结构依赖关系。大量下游任务实验表明,GaussFusion显著提升了高斯表示的迁移能力。特别地,在ModelNet40和ScanObjectNN(PB-T50-RS)上,分别比Gaussian-MAE提升0.61%和3.85%。

原文摘要 · Abstract (English)

3D Gaussian Splatting provides an explicit representation that jointly models geometry and appearance, serving as a scalable foundation for 3D representation learning. Existing pre-training methods for Gaussian representations, such as masked Gaussian reconstruction, primarily capture local structures but offer limited semantic supervision. In this paper, we propose GaussFusion, a multimodal pre-training framework for 3D Gaussian representations. GaussFusion integrates image and text supervision into masked Gaussian modeling through cross-modal semantic alignment, enabling the Gaussian encoder to learn both visual and language-level semantic information during pre-training. To better adapt masked modeling to the non-uniform distribution of Gaussian primitives, we further propose Gaussian Salience-guided Multi-scale Hole Masking (GSHM). GSHM constructs spatially continuous masked regions based on Gaussian salience. By applying hole masks at multiple scales, GSHM encourages the encoder to capture both fine-grained local patterns and broader structural dependencies. Extensive experiments on downstream tasks demonstrate that GaussFusion improves the transferability of Gaussian representations. Notably, GaussFusion outperforms Gaussian-MAE on ModelNet40 and ScanObjectNN (PB-T50-RS) by 0.61\% and 3.85\%, respectively.

3D生成多模态高斯点云预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。