arXiv:2412.08907cs.CV2024-12被引 11

让AI看图定位全球,还能解释推理过程并支持用户互动修正。

Towards Interactive Global Geolocation Assistant

  • 基于视觉语言模型,结合图像线索与世界知识进行定位
  • 在500万图文对数据上训练,国家级精度提升4.57%,城市级提升2.92%
  • 支持用户交互纠错或补充线索,适合需要可解释定位的场景

全球地理定位旨在预测任意位置拍摄图像的地理位置,是计算机视觉中最具挑战性的任务之一。本文提出一种创新的交互式全球定位助手GaGA,基于蓬勃发展的大视觉语言模型(LVLMs)。GaGA从图像中挖掘地理线索,并融合LVLM中蕴含的广泛世界知识,实现地理位置判断,同时提供预测结果的解释与理由。我们设计了一种新型交互式定位方法,超越传统静态推理方式,允许用户干预、纠正或提供线索,使模型更灵活实用。GaGA的开发依赖于新提出的多模态全球定位(MG-Geo)数据集,包含500万高质量图像-文本对。在GWS15k数据集上,GaGA达到当前最佳性能,国家层级准确率提升4.57%,城市层级提升2.92%,树立新基准。这些进展标志着高精度、可交互、全球适用的定位系统的重要突破。

原文摘要 · Abstract (English)

Global geolocation, which seeks to predict the geographical location of images captured anywhere in the world, is one of the most challenging tasks in the field of computer vision. In this paper, we introduce an innovative interactive global geolocation assistant named GaGA, built upon the flourishing large vision-language models (LVLMs). GaGA uncovers geographical clues within images and combines them with the extensive world knowledge embedded in LVLMs to determine the geolocations while also providing justifications and explanations for the prediction results. We further designed a novel interactive geolocation method that surpasses traditional static inference approaches. It allows users to intervene, correct, or provide clues for the predictions, making the model more flexible and practical. The development of GaGA relies on the newly proposed Multi-modal Global Geolocation (MG-Geo) dataset, a comprehensive collection of 5 million high-quality image-text pairs. GaGA achieves state-of-the-art performance on the GWS15k dataset, improving accuracy by 4.57% at the country level and 2.92% at the city level, setting a new benchmark. These advancements represent a significant leap forward in developing highly accurate, interactive geolocation systems with global applicability.

地理定位视觉语言模型交互式系统多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。