用多模态融合提升无人机与地面视角的定位匹配精度。
MAG-VLAQ: Multi-modal Aerial-Ground Query Aggregation for Cross-View Place Recognition

- 结合视觉、激光雷达数据,动态聚合查询特征以对齐跨视角信息。
- 在KITTI360-AG数据集上召回率提升至61.1,接近翻倍。
- 适合做自动驾驶、无人机导航等跨视角定位任务的研究者。
由于地面观测与航空参考之间存在严重的视角、模态和空间结构差异,多模态跨视角场景识别仍是计算机视觉与机器人领域的核心挑战。为此,我们提出MAG-VLAQ——一种基于基础模型增强的查询聚合框架,用于多模态空中-地面跨视角场景识别。具体地,利用预训练基础模型从地面与航空图像中提取密集视觉令牌,并从地面激光雷达数据中提取具表达力的几何令牌,将这些异构令牌映射到共享嵌入空间实现跨模态对齐与融合。作为主要贡献,我们提出基于常微分方程(ODE)的VLAQ机制,将基于ODE的RGB-LiDAR融合与局部聚合查询向量(VLAQ)紧密结合。该设计使查询中心根据融合后的多模态状态动态调整,使最终全局描述符既保留全局检索原型,又对场景特异性视觉与几何证据保持响应,显著提升空中-地面匹配性能。在KITTI360-AG与nuScenes-AG上的大量实验验证了MAG-VLAQ的有效性。特别地,在KITTI360-AG数据集上,我们的方法将最先进性能几乎翻倍,卫星设置下达到61.1 Recall@1,优于对比方法的34.5。
原文摘要 · Abstract (English)
Multi-modal cross-view place recognition remains a fundamental challenge in computer vision and robotics due to the severe viewpoint, modality, and spatial-structure discrepancies between ground observations and aerial references. To address this challenge, we present MAG-VLAQ, a foundation-model-enhanced query aggregation framework for multi-modal aerial-ground cross-view place recognition. Specifically, our approach leverages pre-trained foundation models to extract dense visual tokens from both ground and aerial images, as well as expressive geometric tokens from ground LiDAR observations. These heterogeneous tokens are then projected into a shared embedding space for cross-modal alignment and fusion. As our main contribution, we propose ODE-conditioned VLAQ, which tightly couples neural ordinary differential equations (ODE)-based RGB-LiDAR fusion with vectors of locally aggregated queries (VLAQ). In this design, the VLAQ query centers are dynamically adapted according to the fused multi-modal state. This mechanism allows the final global descriptor to preserve globally learned retrieval prototypes while remaining responsive to scene-specific visual and geometric evidence, significantly improving aerial-ground matching. Extensive experiments on KITTI360-AG and nuScenes-AG validate the effectiveness of our proposed MAG-VLAQ. Notably, on KITTI360-AG, our MAG-VLAQ nearly doubles the state-of-the-art performance, achieving 61.1 Recall@1 in the satellite setting, compared with 34.5 from the closest competing approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。