arXiv:2509.01341cs.CVcs.AI2025-09被引 1

用大模型+检索增强实现高精度街景定位,无需费时训练。

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation

  • 结合多模态大模型与检索增强生成,用相似/不相似位置信息增强图像输入。
  • 在IM2GPS、IM2GPS3k、YFCC4k三个数据集上达到当前最佳准确率。
  • 无需微调,可无缝扩展新数据,适合追求高效部署的地理智能应用。

从图像进行街景级地理定位对导航、位置推荐和城市规划等应用至关重要。随着社交媒体数据和手机摄像头普及,传统计算机视觉方法面临挑战但价值更高。本文提出一种新方法:利用公开的多模态大语言模型,并结合基于SigLIP编码器构建的向量数据库(基于EMP-16和OSV-5M两个大规模数据集)。查询图像通过包含从数据库中检索出的相似与不相似地理位置信息的提示进行增强后,再由多模态大模型处理。该方法在三个主流基准数据集(IM2GPS、IM2GPS3k、YFCC4k)上表现优异,优于现有方法。关键优势在于无需昂贵微调或重新训练,且能轻松扩展新数据源。结果表明,基于检索增强生成的多模态大模型为地理定位提供了替代传统从头训练的新路径,推动更可及、可扩展的GeoAI解决方案发展。

原文摘要 · Abstract (English)

Street-level geolocalization from images is crucial for a wide range of essential applications and services, such as navigation, location-based recommendations, and urban planning. With the growing popularity of social media data and cameras embedded in smartphones, applying traditional computer vision techniques to localize images has become increasingly challenging, yet highly valuable. This paper introduces a novel approach that integrates open-weight and publicly accessible multimodal large language models with retrieval-augmented generation. The method constructs a vector database using the SigLIP encoder on two large-scale datasets (EMP-16 and OSV-5M). Query images are augmented with prompts containing both similar and dissimilar geolocation information retrieved from this database before being processed by the multimodal large language models. Our approach has demonstrated state-of-the-art performance, achieving higher accuracy compared against three widely used benchmark datasets (IM2GPS, IM2GPS3k, and YFCC4k). Importantly, our solution eliminates the need for expensive fine-tuning or retraining and scales seamlessly to incorporate new data sources. The effectiveness of retrieval-augmented generation-based multimodal large language models in geolocation estimation demonstrated by this paper suggests an alternative path to the traditional methods which rely on the training models from scratch, opening new possibilities for more accessible and scalable solutions in GeoAI.

地理定位多模态大模型检索增强GeoAI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。