arXiv:2606.04133cs.CV2026-06

融合网络图片与街景图,精准定位照片拍摄地

Pinpoint: Grounded Worldwide Image Geolocation via Cross-Source Retrieval and Reranking

论文配图:Pinpoint: Grounded Worldwide Image Geolocation via Cross-Source Retrieval and Reranking
图 1 · 摘自论文原文
  • 用对比学习统一训练网络图与街景图的图像-地理嵌入空间
  • 在IM2GPS3k、YFCC4k等数据集上达到最优准确率
  • 无需大模型,推理快且可复现,适合实际地理定位应用

图像地理定位旨在仅通过视觉内容推断照片的拍摄位置。在全球尺度下,该任务仍具挑战性,因视觉线索常模糊、多样且分布不均。以往工作通常将普通网络照片与街景图像的定位视为独立任务,尽管二者各具优势:网络照片更贴近用户自拍图像的外观分布,而街景图像提供更密集的地理覆盖。本文提出Pinpoint,一种基于检索与重排序的粗到精框架,联合利用两类数据源。一个对比学习的图像-地理嵌入器在Flickr上传照片与街景图像上共同训练,学习共享的图像-地理嵌入空间以检索候选位置。随后,基于注意力机制的重排序器结合候选位置的视觉与地理特征,并融合邻近位置的跨源证据,实现定位结果的精确校准。与近期方法不同,Pinpoint不依赖多模态大语言模型,从而提升推理速度与可复现性。在标准基准测试(包括IM2GPS3k、YFCC4k和OSV-5M)上,Pinpoint在所有指标上均达到当前最佳表现。

原文摘要 · Abstract (English)

Image geolocation aims to estimate where a photograph was taken from its visual content. At worldwide scale, this remains challenging because visual evidence is often ambiguous, diverse, and unevenly distributed. Prior work has typically treated geolocation of ordinary internet photos and street-view imagery as separate tasks, despite their complementary strengths: internet photos better match the appearance distribution of user-captured queries, while street-view imagery provides denser, geographically grounded coverage. We present Pinpoint, a retrieve-and-rerank architecture that combines both sources in a coarse-to-fine pipeline. A contrastive image-GPS embedder is trained on both user-uploaded Flickr photos and street-view imagery, learning a shared image-GPS embedding space that is used to retrieve candidate locations. An attention-based reranker then rescores retrieved candidates by combining candidate-level visual and GPS features with cross-source evidence from nearby locations to ground the prediction. Unlike recent prior work, Pinpoint does not rely on multimodal large-language models, making inference faster and more reproducible. Pinpoint achieves state-of-the-art results across all metrics on standard benchmarks for internet photos (IM2GPS3k and YFCC4k) and street-view imagery (OSV-5M).

图像定位跨源检索地理信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。