arXiv:2604.16248cs.CV2026-04中稿 · CVPR

评测视觉语言模型在国家级图像定位中的表现,发现其语义推理有潜力但细节捕捉不足。

Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization

论文配图:Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization
图 1 · 摘自论文原文
  • 用提示词零样本预测图片所属国家,不依赖标注数据或地理元信息。
  • 多模型在三个地理多样数据集上表现差异大,粗略定位有效但精细识别弱。
  • 首次系统对比现代VLM在地理推理任务中的能力,适合研究多模态与地理认知交叉者。

图像地理定位传统上依赖检索式场景识别或基于几何的视觉定位流程。近年来,视觉语言模型(VLMs)在跨模态任务中展现出强大的零样本推理能力,但在地理推断方面的表现仍缺乏深入研究。本文针对国家级别的图像地理定位,仅使用地面视角图像,对多个先进VLM进行系统评估。不依赖图像匹配、GPS元数据或任务特定训练,采用提示词驱动的零样本国家预测方法。所选模型在三个地理分布多样化的数据集上测试,以评估其鲁棒性与泛化能力。结果表明,各模型表现存在显著差异,揭示了语义推理在粗粒度地理定位中的潜力,以及当前VLM在捕捉细微地理线索方面的局限性。本研究首次聚焦于现代VLM在国家级地理定位中的比较,为多模态推理与地理理解交叉领域奠定了基础。

原文摘要 · Abstract (English)

Image geolocalization has traditionally been addressed through retrieval-based place recognition or geometry-based visual localization pipelines. Recent advances in Vision-Language Models (VLMs) have demonstrated strong zero-shot reasoning capabilities across multimodal tasks, yet their performance in geographic inference remains underexplored. In this work, we present a systematic evaluation of multiple state-of-the-art VLMs for country-level image geolocalization using ground-view imagery only. Instead of relying on image matching, GPS metadata, or task-specific training, we evaluate prompt-based country prediction in a zero-shot setting. The selected models are tested on three geographically diverse datasets to assess their robustness and generalization ability. Our results reveal substantial variation across models, highlighting the potential of semantic reasoning for coarse geolocalization and the limitations of current VLMs in capturing fine-grained geographic cues. This study provides the first focused comparison of modern VLMs for country-level geolocalization and establishes a foundation for future research at the intersection of multimodal reasoning and geographic understanding.

视觉语言模型地理定位零样本推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。