arXiv:2512.10419cs.CV2025-12

用跨模态注意力融合俯视图与地面激光,实现精准车辆定位

TransLocNet: Cross-Modal Attention for Aerial-Ground Vehicle Localization with Contrastive Learning

  • 通过双向注意力对齐俯视影像与地面激光数据
  • 在CARLA和KITTI上定位误差降低63%,达亚米级精度
  • 适合自动驾驶中复杂视角下的定位任务

地表与航拍定位困难源于视角和模态差异。我们提出TransLocNet,一种融合激光几何与航拍语义的跨模态注意力框架。将激光扫描投影为鸟瞰表示,并通过双向注意力与航拍特征对齐,再经似然图解码器输出位置与朝向的概率分布。对比学习模块强制共享嵌入空间以增强跨模态对齐。在CARLA和KITTI上的实验表明,TransLocNet优于现有基线,定位误差最高降低63%,实现亚米级、亚度级精度。结果证明该方法在合成与真实场景下均具鲁棒性和泛化能力。

原文摘要 · Abstract (English)

Aerial-ground localization is difficult due to large viewpoint and modality gaps between ground-level LiDAR and overhead imagery. We propose TransLocNet, a cross-modal attention framework that fuses LiDAR geometry with aerial semantic context. LiDAR scans are projected into a bird's-eye-view representation and aligned with aerial features through bidirectional attention, followed by a likelihood map decoder that outputs spatial probability distributions over position and orientation. A contrastive learning module enforces a shared embedding space to improve cross-modal alignment. Experiments on CARLA and KITTI show that TransLocNet outperforms state-of-the-art baselines, reducing localization error by up to 63% and achieving sub-meter, sub-degree accuracy. These results demonstrate that TransLocNet provides robust and generalizable aerial-ground localization in both synthetic and real-world settings.

跨模态定位激光雷达注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。