arXiv:2510.26795cs.CVcs.LG2025-10NeurIPS被引 11

提出混合方法实现大陆级图像精确定位,定位精度超200米。

Scaling Image Geo-Localization to Continent Level

  • 用代理分类任务学习带地理信息的特征表示
  • 在覆盖欧洲大部分区域的数据集上,68%查询定位精度超200米
  • 适合需要跨国家高精度地理定位的研究与应用

在全局范围内精确确定图像地理位置仍是未解难题。传统图像检索方法因图像数量庞大(超1亿)且覆盖不足而效率低下。现有可扩展方案存在权衡:全球分类通常结果粗糙(超过10公里),而地-空视角间的跨视图检索受域差距影响,且多限于小范围研究。本文提出一种混合方法,在大陆尺度实现细粒度地理定位。训练中引入代理分类任务,学习蕴含精确位置信息的丰富特征表示,并结合航空影像嵌入,增强对地面数据稀疏性的鲁棒性。该方法支持直接在覆盖多个国家的区域进行细粒度检索。大量实验表明,本方法在覆盖欧洲大部分区域的数据集上,超过68%的查询可实现200米以内定位。代码已公开于https://scaling-geoloc.github.io。

原文摘要 · Abstract (English)

Determining the precise geographic location of an image at a global scale remains an unsolved challenge. Standard image retrieval techniques are inefficient due to the sheer volume of images (>100M) and fail when coverage is insufficient. Scalable solutions, however, involve a trade-off: global classification typically yields coarse results (10+ kilometers), while cross-view retrieval between ground and aerial imagery suffers from a domain gap and has been primarily studied on smaller regions. This paper introduces a hybrid approach that achieves fine-grained geo-localization across a large geographic expanse the size of a continent. We leverage a proxy classification task during training to learn rich feature representations that implicitly encode precise location information. We combine these learned prototypes with embeddings of aerial imagery to increase robustness to the sparsity of ground-level data. This enables direct, fine-grained retrieval over areas spanning multiple countries. Our extensive evaluation demonstrates that our approach can localize within 200m more than 68\% of queries of a dataset covering a large part of Europe. The code is publicly available at https://scaling-geoloc.github.io.

地理定位图像检索跨模态大规模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。