arXiv:2508.01608cs.CV2025-08被引 5

评测大模型地理定位能力,发现其在高资源地区表现更好。

From Pixels to Places: A Systematic Benchmark for Evaluating Image Geolocalization Ability in Large Language Models

  • 构建多源数据集,系统评估大模型地理定位准确率与推理过程。
  • 闭源模型整体表现优于开源模型,且在北美西欧等地误差更小。
  • 揭示模型存在地理偏见,对欠发达地区识别能力显著下降。

图像地理定位是识别图像中地理位置的重要任务,广泛应用于危机响应、数字取证和位置情报等领域。尽管大语言模型(LLMs)在视觉推理方面取得进展,但其地理定位能力仍缺乏系统评估。本文提出IMAGEO-Bench基准,涵盖全球街景、美国兴趣点(POIs)及私有未见图像三类数据集,系统评估准确率、距离误差、地理空间偏差及推理过程。在10个前沿大模型(含开源自闭源)上实验表明,闭源模型普遍具备更强推理能力;重要的是,模型在高资源区域(如北美、西欧、加州)表现更优,而在欠代表地区性能显著下降。回归分析显示,成功定位主要依赖于城市环境、户外场景、街景图像及可识别地标识别。该基准为理解大模型空间推理能力提供了严谨视角,并为构建地理感知智能系统提供指导。

原文摘要 · Abstract (English)

Image geolocalization, the task of identifying the geographic location depicted in an image, is important for applications in crisis response, digital forensics, and location-based intelligence. While recent advances in large language models (LLMs) offer new opportunities for visual reasoning, their ability to perform image geolocalization remains underexplored. In this study, we introduce a benchmark called IMAGEO-Bench that systematically evaluates accuracy, distance error, geospatial bias, and reasoning process. Our benchmark includes three diverse datasets covering global street scenes, points of interest (POIs) in the United States, and a private collection of unseen images. Through experiments on 10 state-of-the-art LLMs, including both open- and closed-source models, we reveal clear performance disparities, with closed-source models generally showing stronger reasoning. Importantly, we uncover geospatial biases as LLMs tend to perform better in high-resource regions (e.g., North America, Western Europe, and California) while exhibiting degraded performance in underrepresented areas. Regression diagnostics demonstrate that successful geolocalization is primarily dependent on recognizing urban settings, outdoor environments, street-level imagery, and identifiable landmarks. Overall, IMAGEO-Bench provides a rigorous lens into the spatial reasoning capabilities of LLMs and offers implications for building geolocation-aware AI systems.

地理定位大模型偏见分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。