融合语义信息提升机器人在复杂环境下的跨模态定位精度
Semantic-Enhanced Cross-Modal Place Recognition for Robust Robot Localization
- 用VMamba提取图像特征,结合分割掩码增强语义感知
- 在KITTI和KITTI-360上达到当前最优,尤其在视角变化下表现稳健
- 适合需要高鲁棒性定位的自动驾驶与无人巡检场景
在无GPS环境下实现机器人精准定位是一项挑战。基于RGB的视觉位置识别方法易受光照、天气和季节变化影响。现有跨模态方法虽利用RGB图像与3D LiDAR地图的几何特性降低敏感性,但在复杂场景、细粒度或高分辨率匹配以及视角变化时仍表现不足。本文提出语义增强型跨模态位置识别框架SCM-PR,通过引入VMamba骨干网络提取RGB图像特征,设计语义感知特征融合(SAFF)模块整合位置描述符与分割掩码,构建融合语义与几何的LiDAR描述符,并在NetVLAD中加入跨模态语义注意力机制以提升匹配效果。同时,在对比学习框架下设计多视角语义-几何匹配与语义一致性损失。在KITTI与KITTI-360数据集上的实验表明,SCM-PR在跨模态位置识别任务中优于现有方法。
原文摘要 · Abstract (English)
Ensuring accurate localization of robots in environments without GPS capability is a challenging task. Visual Place Recognition (VPR) techniques can potentially achieve this goal, but existing RGB-based methods are sensitive to changes in illumination, weather, and other seasonal changes. Existing cross-modal localization methods leverage the geometric properties of RGB images and 3D LiDAR maps to reduce the sensitivity issues highlighted above. Currently, state-of-the-art methods struggle in complex scenes, fine-grained or high-resolution matching, and situations where changes can occur in viewpoint. In this work, we introduce a framework we call Semantic-Enhanced Cross-Modal Place Recognition (SCM-PR) that combines high-level semantics utilizing RGB images for robust localization in LiDAR maps. Our proposed method introduces: a VMamba backbone for feature extraction of RGB images; a Semantic-Aware Feature Fusion (SAFF) module for using both place descriptors and segmentation masks; LiDAR descriptors that incorporate both semantics and geometry; and a cross-modal semantic attention mechanism in NetVLAD to improve matching. Incorporating the semantic information also was instrumental in designing a Multi-View Semantic-Geometric Matching and a Semantic Consistency Loss, both in a contrastive learning framework. Our experimental work on the KITTI and KITTI-360 datasets show that SCM-PR achieves state-of-the-art performance compared to other cross-modal place recognition methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。