首个无需深度图的单目语义建图系统,支持开放世界语义理解。
OpenMonoGS-SLAM: Monocular Gaussian Splatting SLAM with Open-set Semantics
- 融合3D高斯泼溅与视觉基础模型实现单目建图
- 在封闭与开放语义任务中性能超越现有方法
- 适合无人车、AR/VR等开放环境智能感知应用
同时定位与地图构建(SLAM)是机器人、增强现实/虚拟现实及自动驾驶系统的核心。近年来,空间人工智能的发展推动了将语义理解融入SLAM的需求。现有工作多依赖深度传感器或封闭集语义模型,难以适应开放世界。本文提出OpenMonoGS-SLAM,首个结合3D高斯泼溅与开放词汇语义理解的单目SLAM框架。我们利用MASt3R进行视觉几何重建,SAM与CLIP实现开放词汇语义理解,通过自监督学习目标实现无深度输入、无3D语义真值下的端到端建图。设计了专用于管理高维语义特征的记忆机制,有效构建高斯语义特征图。实验表明,该方法在封闭集与开放集语义分割任务中均达到或超越现有基线性能,且不依赖深度图或语义标注。
原文摘要 · Abstract (English)
Simultaneous Localization and Mapping (SLAM) is a foundational component in robotics, AR/VR, and autonomous systems. With the rising focus on spatial AI in recent years, combining SLAM with semantic understanding has become increasingly important for enabling intelligent perception and interaction. Recent efforts have explored this integration, but they often rely on depth sensors or closed-set semantic models, limiting their scalability and adaptability in open-world environments. In this work, we present OpenMonoGS-SLAM, the first monocular SLAM framework that unifies 3D Gaussian Splatting (3DGS) with open-set semantic understanding. To achieve our goal, we leverage recent advances in Visual Foundation Models (VFMs), including MASt3R for visual geometry and SAM and CLIP for open-vocabulary semantics. These models provide robust generalization across diverse tasks, enabling accurate monocular camera tracking and mapping, as well as a rich understanding of semantics in open-world environments. Our method operates without any depth input or 3D semantic ground truth, relying solely on self-supervised learning objectives. Furthermore, we propose a memory mechanism specifically designed to manage high-dimensional semantic features, which effectively constructs Gaussian semantic feature maps, leading to strong overall performance. Experimental results demonstrate that our approach achieves performance comparable to or surpassing existing baselines in both closed-set and open-set segmentation tasks, all without relying on supplementary sensors such as depth maps or semantic annotations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。