arXiv:2510.20549cs.CVcs.RO2025-10被引 1

用深度学习提升视觉SLAM在弱纹理等复杂场景下的定位精度,助力视障者导航

Deep Learning-Powered Visual SLAM Aimed at Assisting Visually Impaired Navigation

  • 融合SuperPoint与LightGlue进行鲁棒特征提取与匹配
  • 在低纹理、快速运动场景下比ORB-SLAM3平均提升87.84%性能
  • 适用于视障导航等对稳定性要求高的实际应用场景

尽管SLAM技术不断进步,但在低纹理、运动模糊或光照困难等挑战性条件下仍难以稳定运行,这类情况在视障者辅助导航中尤为常见。这些条件会降低定位精度和追踪稳定性,影响导航的可靠性与安全性。为此,本文提出SELM-SLAM3——一种结合SuperPoint与LightGlue的深度学习增强型视觉SLAM框架,实现鲁棒特征提取与匹配。我们在TUM RGB-D、ICL-NUIM和TartanAir数据集上进行了评估,涵盖多样且具有挑战性的场景。实验表明,SELM-SLAM3在平均性能上较传统ORB-SLAM3提升87.84%,超过当前最先进的RGB-D SLAM系统36.77%。该框架在低纹理场景和高速运动条件下表现优异,为视障者导航辅助系统开发提供了可靠平台。

原文摘要 · Abstract (English)

Despite advancements in SLAM technologies, robust operation under challenging conditions such as low-texture, motion-blur, or challenging lighting remains an open challenge. Such conditions are common in applications such as assistive navigation for the visually impaired. These challenges undermine localization accuracy and tracking stability, reducing navigation reliability and safety. To overcome these limitations, we present SELM-SLAM3, a deep learning-enhanced visual SLAM framework that integrates SuperPoint and LightGlue for robust feature extraction and matching. We evaluated our framework using TUM RGB-D, ICL-NUIM, and TartanAir datasets, which feature diverse and challenging scenarios. SELM-SLAM3 outperforms conventional ORB-SLAM3 by an average of 87.84% and exceeds state-of-the-art RGB-D SLAM systems by 36.77%. Our framework demonstrates enhanced performance under challenging conditions, such as low-texture scenes and fast motion, providing a reliable platform for developing navigation aids for the visually impaired.

视觉SLAM视障导航深度学习特征匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。