arXiv:2608.29003cs.CVcs.AI2026-08中稿 · IEEE/RSJ INTERNATI…

用单目视频实现动态场景下鲁棒的语义感知重建

RoSe-SLAM: Robust Semantic-Aware Gaussian Splatting SLAM from Dynamic Monocular Videos

论文配图:RoSe-SLAM: Robust Semantic-Aware Gaussian Splatting SLAM from Dynamic Monocular Videos
图 1 · 摘自论文原文
  • 融合2D基础模型语义特征,增强动态物体识别能力
  • 通过时空运动掩码分离动态与静态成分,提升重建精度
  • 适合长期动态室内环境下的高精度定位与建图

在动态非结构化环境中,传统SLAM系统因假设场景静态而性能显著下降。本文提出鲁棒语义感知高斯点云SLAM(RoSe-SLAM),从未标定单目输入中实现整体语义场景理解,达成精准相机追踪与高质量几何重建。不同于依赖人工标注语义标签的传统方法,本工作利用2D基础模型提取的语义特征,增强动态跟踪与建图性能。通过将丰富语义特征注入高斯场,有效识别动态干扰物并实现多视角语义一致性,显著提升几何重建与场景补全效果。具体地,提出时空运动掩码生成模块,实现长期运动监控与短期瞬态动态捕捉,实现动态物体与静态背景的鲁棒解耦。全局束调整阶段,引入遮挡感知关键帧选择机制,以遮挡程度为指标选取关键帧,并设计多视图语义一致性模块,提升动态环境下的映射质量。结合几何运动线索与语义先验,系统动态过滤不可靠观测,重建准确静态场景几何。在动态TUM、Bonn和Wild-Mocap等基准数据集上的大量实验表明,该方法在轨迹估计与静态场景建图方面均表现优异,优于现有动态RGB SLAM基线,在长时动态室内环境下优势明显。

原文摘要 · Abstract (English)

In dynamic and unstructured environments, conventional SLAM systems generally suffer from significant accuracy degeneration due to their static assumptions. In this work, we propose Robust Semantic-aware Gaussian Splatting SLAM (RoSe-SLAM), to address the dynamic challenge by a holistic semantic scene understanding from uncalibrated monocular inputs, achieving accurate camera tracking and high-quality geometry reconstruction. Unlike conventional semantic SLAM using handcrafted semantic labels, our RoSe-SLAM exploits the semantic feature from 2D foundation model to enhance the dynamic tracking and mapping performance. By distilling the rich semantic features to our Gaussian fields, our method effectively identifies dynamic distractors and achieves semantic-aware multi-view consistency, significantly enhancing the geometric reconstruction and scene inpainting. Specifically, we propose a spatial-temporal motion mask generation module, enabling both long-term motion monitoring and short-term transient dynamics capturing, achieving robust and effective disentanglement of dynamic objects and static backgrounds. During global bundle adjustment, we propose an occlusion-aware keyframe selection mechanism to prioritize the occlusion as metric to pick the keyframes, and a multi-view semantic consistency module to improve the mapping quality in dynamic environments. By combining geometric motion cues with semantic priors, our system dynamically filters unreliable observations and reconstructs accurate static scene geometry. Extensive experiments conducted on benchmark datasets including dynamic TUM, Bonn and Wild-Mocap datasets, demonstrate that our method achieves superior performance in both trajectory estimation and static scene mapping, outperforming existing dynamic RGB SLAM baselines in long-term dynamic indoor environments.

SLAM语义感知高斯点云动态场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。