arXiv:2410.15869cs.ROcs.SY2024-10被引 14

利用环境中的文字线索提升机器人在重复场景下的回环检测能力

Robust Loop Closure by Textual Cues in Challenging Environments

  • 通过OCR提取场景文字,结合激光里程计构建文本地图
  • 在走廊等重复环境中,回环检测准确率显著优于纯视觉/激光方法
  • 适合需要高鲁棒性导航的仓储、隧道等复杂场景

回环检测是机器人导航中的关键任务。然而,现有方法大多依赖环境的隐式或启发式特征,在走廊、隧道和仓库等特征缺失、退化且重复(FDR)的环境中仍易失效。事实上,即使对人类而言,这类环境也极具挑战性,而环境中显式的文字线索往往是最有效的辅助。受此启发,我们提出一种基于显式可读文字线索的多模态回环检测方法。具体而言,该方法首先通过光学字符识别(OCR)提取场景中的文字实体,再结合高精度激光里程计构建文本局部地图,并最终采用图论方法识别回环事件。实验结果表明,该方法在依赖单一视觉或激光传感器的方法上表现更优。为促进社区发展,我们已公开源代码与数据集:https://github.com/TongxingJin/TXTLCD。

原文摘要 · Abstract (English)

Loop closure is an important task in robot navigation. However, existing methods mostly rely on some implicit or heuristic features of the environment, which can still fail to work in common environments such as corridors, tunnels, and warehouses. Indeed, navigating in such featureless, degenerative, and repetitive (FDR) environments would also pose a significant challenge even for humans, but explicit text cues in the surroundings often provide the best assistance. This inspires us to propose a multi-modal loop closure method based on explicit human-readable textual cues in FDR environments. Specifically, our approach first extracts scene text entities based on Optical Character Recognition (OCR), then creates a local map of text cues based on accurate LiDAR odometry and finally identifies loop closure events by a graph-theoretic scheme. Experiment results demonstrate that this approach has superior performance over existing methods that rely solely on visual and LiDAR sensors. To benefit the community, we release the source code and datasets at \url{https://github.com/TongxingJin/TXTLCD}.

回环检测文本线索机器人导航多模态感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。