arXiv:2505.14068cs.CV2025-05综述被引 2

综述多模态地点识别技术,揭示当前挑战与未来方向

Place Recognition Meet Multiple Modalitie: A Comprehensive Review, Current Challenges and Future Directions

  • 梳理CNN、Transformer与跨模态三类主流方法
  • 提出融合视觉、激光与文本提升环境鲁棒性
  • 适合自动驾驶与长期导航研究者参考

地点识别是车辆导航与建图的核心,对判断位置是否已访问至关重要,广泛应用于同时定位与地图构建(SLAM)中的回环检测及复杂环境下的长期导航。本文全面回顾了近期地点识别进展,聚焦三大方法范式:基于卷积神经网络(CNN)的方法、基于Transformer的框架以及跨模态策略。首先阐明其在自主系统中的关键作用;随后梳理CNN方法在大规模场景下鲁棒视觉描述子学习与可扩展性方面的贡献;接着分析基于Transformer模型通过自注意力机制捕捉全局依赖、提升跨场景泛化能力的优势;进一步探讨融合激光雷达、视觉与文本等异构数据的跨模态方法,显著增强对视角、光照与季节变化的鲁棒性。总结了常用数据集与评估指标,并指出当前挑战,包括域适应、实时性能与终身学习,为未来研究提供方向。领先方法的统一框架、代码库及实验结果见 https://github.com/CV4RA/SOTA-Place-Recognitioner。

原文摘要 · Abstract (English)

Place recognition is a cornerstone of vehicle navigation and mapping, which is pivotal in enabling systems to determine whether a location has been previously visited. This capability is critical for tasks such as loop closure in Simultaneous Localization and Mapping (SLAM) and long-term navigation under varying environmental conditions. In this survey, we comprehensively review recent advancements in place recognition, emphasizing three representative methodological paradigms: Convolutional Neural Network (CNN)-based approaches, Transformer-based frameworks, and cross-modal strategies. We begin by elucidating the significance of place recognition within the broader context of autonomous systems. Subsequently, we trace the evolution of CNN-based methods, highlighting their contributions to robust visual descriptor learning and scalability in large-scale environments. We then examine the emerging class of Transformer-based models, which leverage self-attention mechanisms to capture global dependencies and offer improved generalization across diverse scenes. Furthermore, we discuss cross-modal approaches that integrate heterogeneous data sources such as Lidar, vision, and text description, thereby enhancing resilience to viewpoint, illumination, and seasonal variations. We also summarize standard datasets and evaluation metrics widely adopted in the literature. Finally, we identify current research challenges and outline prospective directions, including domain adaptation, real-time performance, and lifelong learning, to inspire future advancements in this domain. The unified framework of leading-edge place recognition methods, i.e., code library, and the results of their experimental evaluations are available at https://github.com/CV4RA/SOTA-Place-Recognitioner.

地点识别多模态SLAM自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。