arXiv:2503.18725cs.CV2025-03CVPR被引 34

通过细粒度特征匹配实现地面与航拍图像的精准定位

FG$^2$: Fine-Grained Cross-View Localization by Fine-Grained Feature Matching

  • 将地面图像特征映射为3D点云,筛选高度维度特征生成鸟瞰图平面
  • 在跨视图点对应关系中采样稀疏匹配,用普鲁克斯特对齐计算姿态
  • 在VIGOR数据集上定位误差降低28%,适合弱监督下的跨视角定位

我们提出一种新型细粒度跨视图定位方法,通过匹配地面图像与航拍图像之间的细粒度特征,估计地面图像在周围环境航拍图像中的3自由度姿态。该方法通过将地面图像特征映射至3D点云,并学习沿高度维度选择特征,将3D点云池化为鸟瞰图(BEV)平面,从而追踪每个地表特征对鸟瞰表示的贡献。随后,从两平面点对应关系中采样一组稀疏匹配,并使用普鲁克斯特对齐算法计算其相对姿态。相比现有最优方法,本方法在VIGOR跨区域测试集上的平均定位误差降低了28%。定性结果表明,该方法通过相机姿态的弱监督学习,实现了地面与航拍视图间的语义一致匹配。

原文摘要 · Abstract (English)

We propose a novel fine-grained cross-view localization method that estimates the 3 Degrees of Freedom pose of a ground-level image in an aerial image of the surroundings by matching fine-grained features between the two images. The pose is estimated by aligning a point plane generated from the ground image with a point plane sampled from the aerial image. To generate the ground points, we first map ground image features to a 3D point cloud. Our method then learns to select features along the height dimension to pool the 3D points to a Bird's-Eye-View (BEV) plane. This selection enables us to trace which feature in the ground image contributes to the BEV representation. Next, we sample a set of sparse matches from computed point correspondences between the two point planes and compute their relative pose using Procrustes alignment. Compared to the previous state-of-the-art, our method reduces the mean localization error by 28% on the VIGOR cross-area test set. Qualitative results show that our method learns semantically consistent matches across ground and aerial views through weakly supervised learning from the camera pose.

跨视图定位细粒度匹配鸟瞰图姿态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。