用两阶段检索提升窄视角街景定位精度
VICI: VLM-Instructed Cross-view Image-localisation

- 先检索候选卫星图嵌入,再精修前几名结果
- 在University-1652数据集上达R@1 98.7%、R@10 100%
- 适合真实场景中低视角、未知参数的定位任务
本文提出VICI方法,针对UAVM 2025挑战赛中的窄视场街景图像与卫星影像匹配问题,使用University-1652数据集。随着全景跨视图地理定位性能接近极限,探索更贴近现实的设定愈发重要:真实场景中街景查询通常为有限视场、相机参数未知的图像。本工作聚焦在该约束下可达到的最高性能,突破现有架构极限。方法采用两阶段策略:首先对给定查询检索候选卫星图像嵌入,随后通过重排序阶段精细提升前几名候选者的匹配精度。该设计有效应对视角与尺度的巨大差异。实验表明,该方法在University-1652数据集上实现R@1 98.7%、R@10 100%的检索率,验证了优化检索与重排序策略在提升实际地理定位性能方面的潜力。代码已开源。
原文摘要 · Abstract (English)
In this paper, we present a high-performing solution to the UAVM 2025 Challenge, which focuses on matching narrow FOV street-level images to corresponding satellite imagery using the University-1652 dataset. As panoramic Cross-View Geo-Localisation nears peak performance, it becomes increasingly important to explore more practical problem formulations. Real-world scenarios rarely offer panoramic street-level queries; instead, queries typically consist of limited-FOV images captured with unknown camera parameters. Our work prioritises discovering the highest achievable performance under these constraints, pushing the limits of existing architectures. Our method begins by retrieving candidate satellite image embeddings for a given query, followed by a re-ranking stage that selectively enhances retrieval accuracy within the top candidates. This two-stage approach enables more precise matching, even under the significant viewpoint and scale variations inherent in the task. Through experimentation, we demonstrate that our approach achieves competitive results -specifically attaining R@1 and R@10 retrieval rates of \topone\% and \topten\% respectively. This underscores the potential of optimised retrieval and re-ranking strategies in advancing practical geo-localisation performance. Code is available at https://github.com/tavisshore/VICI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。