融合语义与3D高斯表示,提升动态场景下定位与建图精度。
STAMICS: Splat, Track And Map with Integrated Consistency and Semantics for Dense RGB-D SLAM
- 用3D高斯表示场景,实现高保真重建。
- 通过图聚类保证时序语义一致性,减少重建误差。
- 支持开放词汇识别未见物体,适合复杂动态环境。
同时定位与建图(SLAM)是机器人自主导航和环境理解的关键任务。现有方法主要依赖几何线索进行建图与定位,但在动态或密集场景中常无法保证语义一致性。为此,我们提出STAMICS,一种将语义信息与3D高斯表示相结合的新方法,以提升定位与建图的准确性。STAMICS包含三个核心组件:基于3D高斯的场景表示,用于高保真重建;基于图的聚类技术,以强制时序语义一致性;以及开放词汇系统,可对未见物体进行分类。大量实验表明,STAMICS显著改善了相机位姿估计与地图质量,在多个指标上超越当前最优方法,同时降低重建误差。代码将开源。
原文摘要 · Abstract (English)
Simultaneous Localization and Mapping (SLAM) is a critical task in robotics, enabling systems to autonomously navigate and understand complex environments. Current SLAM approaches predominantly rely on geometric cues for mapping and localization, but they often fail to ensure semantic consistency, particularly in dynamic or densely populated scenes. To address this limitation, we introduce STAMICS, a novel method that integrates semantic information with 3D Gaussian representations to enhance both localization and mapping accuracy. STAMICS consists of three key components: a 3D Gaussian-based scene representation for high-fidelity reconstruction, a graph-based clustering technique that enforces temporal semantic consistency, and an open-vocabulary system that allows for the classification of unseen objects. Extensive experiments show that STAMICS significantly improves camera pose estimation and map quality, outperforming state-of-the-art methods while reducing reconstruction errors. Code will be public available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。