无需真实位置信息,用自监督方法实现大规模激光雷达定位
S-BEVLoc: BEV-based Self-supervised Framework for Large-scale LiDAR Global Localization
- 基于鸟瞰图构建自监督训练三元组,避免依赖真实姿态标注
- 在KITTI和NCLT数据集上达到顶尖的场景识别与定位性能
- 适合需要低成本、可扩展定位系统的自动驾驶与机器人应用
基于激光雷达的全局定位是同时定位与地图构建(SLAM)的关键组件,有助于回环检测与重定位。现有方法依赖于GPS或SLAM里程计获取的真值位姿进行网络训练,但高精度真值位姿获取成本高昂。本文提出S-BEVLoc,一种基于鸟瞰图(BEV)的自监督框架,无需真值位姿,具备高度可扩展性。通过利用关键点周围BEV块间的已知地理距离,从单个BEV图像构建训练三元组。采用卷积神经网络(CNN)提取局部特征,使用NetVLAD聚合全局描述符,并引入SoftCos损失增强三元组学习。在大规模KITTI和NCLT数据集上的实验表明,S-BEVLoc在场景识别、回环检测和全局定位任务中均达到当前最优性能,且相比监督方法更易扩展。
原文摘要 · Abstract (English)
LiDAR-based global localization is an essential component of simultaneous localization and mapping (SLAM), which helps loop closure and re-localization. Current approaches rely on ground-truth poses obtained from GPS or SLAM odometry to supervise network training. Despite the great success of these supervised approaches, substantial cost and effort are required for high-precision ground-truth pose acquisition. In this work, we propose S-BEVLoc, a novel self-supervised framework based on bird's-eye view (BEV) for LiDAR global localization, which eliminates the need for ground-truth poses and is highly scalable. We construct training triplets from single BEV images by leveraging the known geographic distances between keypoint-centered BEV patches. Convolutional neural network (CNN) is used to extract local features, and NetVLAD is employed to aggregate global descriptors. Moreover, we introduce SoftCos loss to enhance learning from the generated triplets. Experimental results on the large-scale KITTI and NCLT datasets show that S-BEVLoc achieves state-of-the-art performance in place recognition, loop closure, and global localization tasks, while offering scalability that would require extra effort for supervised approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。