arXiv:2604.04513cs.CVcs.RO2026-04

用多视角金字塔变换器融合激光雷达特征,提升复杂环境下的定位识别精度。

MPTF-Net: Multi-view Pyramid Transformer Fusion Network for LiDAR-based Place Recognition

  • 基于NDT的多通道鸟瞰图编码,捕捉局部几何与强度分布细节。
  • 在nuScenes数据集上达96.31%召回率,推理仅耗时10.02毫秒。
  • 适合实时自动驾驶系统,尤其在重复或复杂场景中表现优异。

基于激光雷达的地点识别(LPR)对大规模SLAM系统的全局定位和回环检测至关重要。现有方法通常从距离图像或鸟瞰图(BEV)表示构建全局描述符进行匹配。由于具备显式的二维空间布局编码和高效检索能力,BEV被广泛采用。然而,传统BEV表示依赖简单统计聚合,难以捕捉细粒度几何结构,在复杂或重复环境中性能下降。为此,我们提出MPTF-Net,一种新颖的多视角多尺度金字塔变换器融合网络。核心贡献是基于NDT的多通道BEV编码,通过正态分布变换显式建模局部几何复杂性和强度分布,提供抗噪的结构先验。为有效融合这些特征,我们设计定制化的金字塔变换器模块,捕获距离图像视图(RIV)与NDT-BEV在多空间尺度上的跨视图交互关系。在nuScenes、KITTI和NCLT数据集上的大量实验表明,MPTF-Net达到领先性能,尤其在nuScenes波士顿分割上实现96.31%的Recall@1,同时保持仅10.02毫秒的推理延迟,适用于实时自主无人系统。

原文摘要 · Abstract (English)

LiDAR-based place recognition (LPR) is essential for global localization and loop-closure detection in large-scale SLAM systems. Existing methods typically construct global descriptors from Range Images or BEV representations for matching. BEV is widely adopted due to its explicit 2D spatial layout encoding and efficient retrieval. However, conventional BEV representations rely on simple statistical aggregation, which fails to capture fine-grained geometric structures, leading to performance degradation in complex or repetitive environments. To address this, we propose MPTF-Net, a novel multi-view multi-scale pyramid Transformer fusion network. Our core contribution is a multi-channel NDT-based BEV encoding that explicitly models local geometric complexity and intensity distributions via Normal Distribution Transform, providing a noise-resilient structural prior. To effectively integrate these features, we develop a customized pyramid Transformer module that captures cross-view interactive correlations between Range Image Views (RIV) and NDT-BEV at multiple spatial scales. Extensive experiments on the nuScenes, KITTI and NCLT datasets demonstrate that MPTF-Net achieves state-of-the-art performance, specifically attaining a Recall@1 of 96.31\% on the nuScenes Boston split while maintaining an inference latency of only 10.02 ms, making it highly suitable for real-time autonomous unmanned systems.

激光雷达地点识别变换器自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。