arXiv:2503.16423cs.CVcs.LG2025-03被引 1

让AI能对话式讲解照片位置信息,还支持地理知识问答。

GAEA: A Geolocation Aware Conversational Assistant

  • 用地图数据构建问答对训练对话型定位模型
  • 在3.5千组图像上比最强模型高7.2%准确率
  • 适合需要地理知识交互的智能助手研究者

图像地理定位传统上要求模型预测精确的GPS坐标,但用户无法进一步获取位置背景知识。现有大型多模态模型虽可尝试此任务,但在专业场景下仍表现不佳。为此,本文提出对话式地理定位助手GAEA,可按需提供图像位置的相关信息。由于缺乏大规模训练数据,我们构建了包含80多万张图像和约140万条问答对的GAEA-1.4M数据集,利用OpenStreetMap(OSM)属性与地理上下文线索生成。为评估性能,我们设计了一个包含3500个图像-文本对的多样性基准GAEA-Bench。评估涵盖11个主流开源与闭源模型,结果表明,GAEA相比最佳开源模型LLaVA-OneVision提升18.2%,相比最佳闭源模型GPT-4o提升7.2%。相关数据集、模型与代码均已公开。

原文摘要 · Abstract (English)

Image geolocalization, in which an AI model traditionally predicts the precise GPS coordinates of an image, is a challenging task with many downstream applications. However, the user cannot utilize the model to further their knowledge beyond the GPS coordinates; the model lacks an understanding of the location and the conversational ability to communicate with the user. In recent days, with the tremendous progress of large multimodal models (LMMs) -- proprietary and open-source -- researchers have attempted to geolocalize images via LMMs. However, the issues remain unaddressed; beyond general tasks, for more specialized downstream tasks, such as geolocalization, LMMs struggle. In this work, we propose solving this problem by introducing a conversational model, GAEA, that provides information regarding the location of an image as the user requires. No large-scale dataset enabling the training of such a model exists. Thus, we propose GAEA-1.4M, a comprehensive dataset comprising over 800k images and approximately 1.4M question-answer pairs, constructed by leveraging OpenStreetMap (OSM) attributes and geographical context clues. For quantitative evaluation, we propose a diverse benchmark, GAEA-Bench, comprising 3.5k image-text pairs to evaluate conversational capabilities equipped with diverse question types. We consider 11 state-of-the-art open-source and proprietary LMMs and demonstrate that GAEA significantly outperforms the best open-source model, LLaVA-OneVision, by 18.2% and the best proprietary model, GPT-4o, by 7.2%. Our dataset, model and codes are available.

地理定位对话系统多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。