X-GS让3D高斯点云同时感知与思考,实现实时语义建图和多模态任务。
X-GS: An Extensible Framework for Perceiving and Thinking via 3D Gaussian Splatting
- 统一多种3D高斯方法,支持实时在线语义建图。
- 结合视觉大模型提升建图效率,加速语义融合。
- 兼容对比与生成式多模态模型,可做3D视觉定位与场景描述。
3D高斯点绘(3DGS)已成为新视角合成的强大技术,并扩展至众多空间智能应用。然而,现有大多数3DGS方法各自孤立,仅聚焦特定领域。本文提出X-GS,一个可扩展框架,包含两大组件:X-GS-Perceiver将多种3DGS技术统一,实现实时在线语义建图;X-GS-Thinker集成多模态模型,使其无缝对接感知结果以完成下游任务。在实现中,Perceiver利用最新视觉基础模型提升在线建图性能,并采用三项关键机制加速语义蒸馏。Thinker可基于对比型与生成型视觉-语言模型构建,借助Perceiver的语义高斯点云,实现3D视觉定位与场景描述等能力。多样基准测试表明,该框架兼具高效性与全新多模态功能。
原文摘要 · Abstract (English)
3D Gaussian Splatting (3DGS) has emerged as a powerful technique for novel view synthesis, subsequently extending into numerous spatial AI applications. However, most existing 3DGS methods operate in isolation, focusing on specific domains. In this paper, we introduce X-GS, an extensible framework consisting of two major components. The X-GS-Perceiver unifies a broad range of 3DGS techniques to enable real-time online SLAM with semantic distillation. The X-GS-Thinker accommodates multimodal models, enabling them to seamlessly interface with the Perceiver to complete downstream tasks. In our implementation of X-GS, the Perceiver leverages the latest vision foundation models to improve online SLAM performance and employs three key mechanisms to accelerate semantic distillation. The Thinker can be built upon both contrastive and generative vision-language models and utilizes the Perceiver's semantic Gaussian splats to unlock capabilities such as 3D visual grounding and scene captioning. Experimental results on diverse benchmarks demonstrate the efficiency and newly unlocked multimodal capabilities of the X-GS framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。