用位置注意力提升全球图像定位,让相似画面不再混淆地点。
When Vision Misleads, Let Location Speak: A Worldwide Image Geo-Localization Method via Location Attention Mechanism and Large Multimodal Models

- 引入位置注意力机制,结合大模型增强地理特征区分能力。
- 在多个数据集上实现1公里内定位准确率提升最高达9.75%。
- 适合需要高精度地理定位的导航、地图与跨域图像检索场景。
全球图像地理定位旨在确定图像的拍摄位置。现有方法常因将视觉相似但地理不同的场景错误匹配,导致实际应用可靠性受限。为此,我们提出TransGeoCLIP,一种基于检索的新型框架,融合位置注意力机制与大视觉语言模型(LMMs)。通过带位置注意力的Transformer编码器处理经纬度坐标,有效区分视觉相似图像的地理特征。该框架包含两阶段:1)构建检索数据库,利用带位置注意力的Transformer编码标注的GPS坐标,增强位置语义,并通过CLIP实现图像-文本-地理坐标联合嵌入;2)检索增强推理,借助大模型从检索结果中推断最终位置。在IM2GPS、IM2GPS3k、YFCC4k和YFCC26k等多个数据集上的实验表明,TransGeoCLIP显著提升了视觉相似图像的定位性能。尤其在街级定位(误差小于1公里)上,准确率分别超越当前最优方法1.5%、1.07%、7.18%和9.75%。
原文摘要 · Abstract (English)
Worldwide image geo-localization aims to determine the capture location of an image on a global scale. Existing methods often mislocalize images by matching them to visually similar scenes from different geographic regions, which limits reliability in practical applications. To address this issue, we propose TransGeoCLIP, a novel retrieval-based framework that integrates a location attention mechanism and large multimodal models (LMMs). Using the Transformer encoder with location attention to encode GPS coordinates, TransGeoCLIP can effectively distinguish geographic features among visually similar images. The framework consists of two stages: 1) Retrieval database construction, which employs Transformers equipped with location attention mechanisms to encode labeled GPS coordinates and enhance location semantics, subsequently enables joint image-text-GPS embedding through CLIP; 2) Retrieval-augmented inference, which leverages LMMs to infer the final image location prediction from retrieved database results. Extensive experimental results on diverse datasets, including IM2GPS, IM2GPS3k, YFCC4k, and YFCC26k, demonstrate that TransGeoCLIP significantly enhances localization performance for visually similar images. Particularly, street-level localization accuracy (within 1 km error) is substantially improved, surpassing state-of-the-art methods by 1.5%, 1.07%, 7.18%, and 9.75% on these benchmarks, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。