用噪声定位数据训练卫星与地面图像匹配模型,提升定位精度。
Weakly-supervised Camera Localization by Ground-to-satellite Image Registration

- 仅需噪声定位标签,通过对比学习提取跨视角特征
- 在跨区域测试中优于依赖精确标签的现有方法
- 适合无高精度定位设备但需城市级定位的场景
地面到卫星图像匹配最初用于城市尺度的地面相机定位。本文针对在获得粗略位置和朝向(来自城市级检索或消费级GPS/指南针)后,如何通过地面-卫星图像匹配提升相机位姿精度的问题。现有基于学习的方法需要地面图像的精确GPS标签进行训练,但获取此类标签困难,通常需昂贵的实时动态定位(RTK)设备,且易受信号遮挡、多路径干扰等问题影响。为缓解此问题,本文提出一种弱监督学习策略,在仅使用地面图像噪声姿态标签的情况下进行网络训练。该方法为每张地面图像构建正负卫星图像,并利用对比学习学习对平移估计有用的地面与卫星图像特征表示。同时提出一种自监督策略,用于跨视角图像相对旋转估计,通过生成伪查询与参考图像对来训练网络。实验结果表明,该弱监督学习策略在跨区域评估中表现优于近期依赖精确姿态标签的先进方法。
原文摘要 · Abstract (English)
The ground-to-satellite image matching/retrieval was initially proposed for city-scale ground camera localization. This work addresses the problem of improving camera pose accuracy by ground-to-satellite image matching after a coarse location and orientation have been obtained, either from the city-scale retrieval or from consumer-level GPS and compass sensors. Existing learning-based methods for solving this task require accurate GPS labels of ground images for network training. However, obtaining such accurate GPS labels is difficult, often requiring an expensive {\color{black}Real Time Kinematics (RTK)} setup and suffering from signal occlusion, multi-path signal disruptions, \etc. To alleviate this issue, this paper proposes a weakly supervised learning strategy for ground-to-satellite image registration when only noisy pose labels for ground images are available for network training. It derives positive and negative satellite images for each ground image and leverages contrastive learning to learn feature representations for ground and satellite images useful for translation estimation. We also propose a self-supervision strategy for cross-view image relative rotation estimation, which trains the network by creating pseudo query and reference image pairs. Experimental results show that our weakly supervised learning strategy achieves the best performance on cross-area evaluation compared to recent state-of-the-art methods that are reliant on accurate pose labels for supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。