用稀疏语义地图实现高效精准定位,定位精度提升5倍以上。
SparseLoc: Sparse Open-Set Landmark-based Global Localization for Autonomous Navigation
- 基于视觉-语言模型生成零样本稀疏语义拓扑地图
- 仅用密集地图1/500点数,误差低于5米2度
- 结合蒙特卡洛与后优化策略,适合自动驾驶场景
全局定位是自主导航中的关键问题,可在无GPS环境下实现精确定位。现有技术多依赖密集激光雷达地图,虽精度高但存储与计算开销大。近期方法尝试使用稀疏地图和学习特征,但鲁棒性与泛化能力不足。本文提出SparseLoc,利用视觉-语言基础模型以零样本方式生成稀疏、语义-拓扑地图,并结合改进的蒙特卡洛定位框架及新型后期优化策略,提升位姿估计精度。通过构建紧凑且高度可区分的地图并设计合理优化流程,克服了现有方法局限。系统在定位精度上相比现有稀疏映射方法提升超过5倍;尽管仅使用密集方法1/500的点数,仍能在KITTI序列上保持平均全局定位误差低于5米、2度。
原文摘要 · Abstract (English)
Global localization is a critical problem in autonomous navigation, enabling precise positioning without reliance on GPS. Modern global localization techniques often depend on dense LiDAR maps, which, while precise, require extensive storage and computational resources. Recent approaches have explored alternative methods, such as sparse maps and learned features, but they suffer from poor robustness and generalization. We propose SparseLoc, a global localization framework that leverages vision-language foundation models to generate sparse, semantic-topometric maps in a zero-shot manner. It combines this map representation with a Monte Carlo localization scheme enhanced by a novel late optimization strategy, ensuring improved pose estimation. By constructing compact yet highly discriminative maps and refining localization through a carefully designed optimization schedule, SparseLoc overcomes the limitations of existing techniques, offering a more efficient and robust solution for global localization. Our system achieves over a 5X improvement in localization accuracy compared to existing sparse mapping techniques. Despite utilizing only 1/500th of the points of dense mapping methods, it achieves comparable performance, maintaining an average global localization error below 5m and 2 degrees on KITTI sequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。