通过深度提升的局部特征匹配,实现高精度可解释的跨视角定位。
Loc$^2$: Interpretable Cross-View Localization via Depth-Lifted Local Feature Matching
- 直接学习地面与航拍图像的像素级对应关系,无需像素标注。
- 在跨区域和未知朝向场景下达到当前最优定位精度。
- 结果可解释性强,支持异常值剔除与可视化验证。
我们提出一种精确且可解释的细粒度跨视角定位方法,通过匹配地面图像与参考航拍图像的局部特征,估计地面图像的3自由度(DoF)姿态。不同于依赖全局描述符或鸟瞰图(BEV)变换的先前方法,本方法利用相机位姿进行弱监督,直接学习地面-航拍图像平面间的对应关系。匹配的地面点通过单目深度预测提升至BEV空间,并采用尺度感知的Procrustes对齐估计相机旋转、平移及相对深度与航拍度量空间之间的尺度。该方法轻量、端到端可训练,且无需像素级标注。实验表明,在跨区域测试与未知朝向等挑战性场景中性能领先。此外,该方法具备强可解释性:对应关系质量直接反映定位精度,支持通过RANSAC进行异常值剔除;将重缩放后的地面布局叠加至航拍图像上,可直观呈现定位效果。
原文摘要 · Abstract (English)
We propose an accurate and interpretable fine-grained cross-view localization method that estimates the 3 Degrees of Freedom (DoF) pose of a ground-level image by matching its local features with a reference aerial image. Unlike prior approaches that rely on global descriptors or bird's-eye-view (BEV) transformations, our method directly learns ground-aerial image-plane correspondences using weak supervision from camera poses. The matched ground points are lifted into BEV space with monocular depth predictions, and scale-aware Procrustes alignment is then applied to estimate camera rotation, translation, and optionally the scale between relative depth and the aerial metric space. This formulation is lightweight, end-to-end trainable, and requires no pixel-level annotations. Experiments show state-of-the-art accuracy in challenging scenarios such as cross-area testing and unknown orientation. Furthermore, our method offers strong interpretability: correspondence quality directly reflects localization accuracy and enables outlier rejection via RANSAC, while overlaying the re-scaled ground layout on the aerial image provides an intuitive visual cue of localization performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。