arXiv:2507.00886cs.CVcs.RO2025-07被引 14

用语言对齐的高斯点构建3D视觉语言模型,提升场景理解能力

GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond

  • 将语言特征直接嵌入高斯点,实现早期模态对齐
  • 通过双路稀疏化生成任务相关的全局与局部场景令牌
  • 首个基于高斯点渲染的3D VLM,跨域性能提升5倍

随着多模态语言模型的发展,其在3D场景理解中的应用成为快速拓展的前沿领域,推动了3D视觉语言模型(VLM)的兴起。现有方法严重依赖目标检测器,带来处理瓶颈和分类灵活性不足的问题。为此,我们提出一种面向3D高斯点云场景的场景中心型3D VLM,采用语言与任务感知的场景表征。通过将语言信息关联至每个高斯原语,直接将丰富的语言特征嵌入3D场景表示中,实现早期模态对齐。为处理由此产生的密集表示,我们引入双路稀疏化机制,通过任务引导和位置引导路径,将密集表征提炼为紧凑的任务相关令牌,生成稀疏的、任务感知的全局与局部场景令牌。值得注意的是,我们首次构建基于高斯点渲染的3D VLM,利用标准RGB图像生成的逼真3D表示,在跨域设置下性能较先前3D VLM提升五倍。

原文摘要 · Abstract (English)

As multimodal language models advance, their application to 3D scene understanding is a fast-growing frontier, driving the development of 3D Vision-Language Models (VLMs). Current methods show strong dependence on object detectors, introducing processing bottlenecks and limitations in taxonomic flexibility. To address these limitations, we propose a scene-centric 3D VLM for 3D Gaussian splat scenes that employs language- and task-aware scene representations. Our approach directly embeds rich linguistic features into the 3D scene representation by associating language with each Gaussian primitive, achieving early modality alignment. To process the resulting dense representations, we introduce a dual sparsifier that distills them into compact, task-relevant tokens via task-guided and location-guided pathways, producing sparse, task-aware global and local scene tokens. Notably, we present the first Gaussian splatting-based VLM, leveraging photorealistic 3D representations derived from standard RGB images, demonstrating strong generalization: it improves performance of prior 3D VLM five folds, in out-of-the-domain settings.

3D视觉语言高斯点场景理解多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。