arXiv:2510.02616cs.RO2025-10被引 1

实时动态环境语义视觉定位与建图,精度高且稳定

RSV-SLAM: Toward Real-Time Semantic Visual SLAM in Indoor Dynamic Environments

  • 用深度学习融合语义信息,识别并隔离动态物体
  • 通过扩展卡尔曼滤波检测暂时静止的动态物体,提升鲁棒性
  • 生成网络填补动态物体遮挡区域,适合智能机器人应用

视觉同步定位与建图(SLAM)在服务机器人等领域至关重要。现有视觉SLAM多假设环境静态,在动态场景中表现不佳。本文提出一种面向动态环境的实时语义RGBD SLAM方法,能有效检测移动物体并维护静态地图以保障相机追踪稳定性。核心创新在于将基于深度学习的语义信息融入SLAM系统,减轻动态物体影响;同时引入扩展卡尔曼滤波优化语义分割,识别可能短暂静止的动态物体;还设计生成网络填补动态物体遮挡区域的图像缺失部分。该模块化框架部署于ROS平台,可在GTX1080上实现约22帧/秒的运行速度。在TUM数据集动态序列上的基准测试表明,本方法在接近实时条件下定位误差优于或媲美当前先进方法,源代码已公开。

原文摘要 · Abstract (English)

Simultaneous Localization and Mapping (SLAM) plays an important role in many robotics fields, including social robots. Many of the available visual SLAM methods are based on the assumption of a static world and struggle in dynamic environments. In the current study, we introduce a real-time semantic RGBD SLAM approach designed specifically for dynamic environments. Our proposed system can effectively detect moving objects and maintain a static map to ensure robust camera tracking. The key innovation of our approach is the incorporation of deep learning-based semantic information into SLAM systems to mitigate the impact of dynamic objects. Additionally, we enhance the semantic segmentation process by integrating an Extended Kalman filter to identify dynamic objects that may be temporarily idle. We have also implemented a generative network to fill in the missing regions of input images belonging to dynamic objects. This highly modular framework has been implemented on the ROS platform and can achieve around 22 fps on a GTX1080. Benchmarking the developed pipeline on dynamic sequences from the TUM dataset suggests that the proposed approach delivers competitive localization error in comparison with the state-of-the-art methods, all while operating in near real-time. The source code is publicly available.

SLAM语义感知动态环境实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。