首个专家级地理推理链数据集,揭示大模型解释能力远不如人类。
GeoRC: A Benchmark for Geolocation Reasoning Chains
- 基于顶尖玩家的实测推理链构建基准,覆盖500张地图场景
- 大模型定位准但解释差,小模型几乎全靠编造
- 适合研究模型可解释性与细粒度视觉理解的学者
视觉语言模型在识别照片全局位置方面表现优异,其准确率接近顶尖人类专家。然而,许多模型在解释预测依据时表现极差,即使位置判断正确亦然。本文提出GeoRC,首个源自冠军级GeoGuessr玩家(含世界冠军)的地理推理链基准,包含800条“真实”推理链,覆盖500个查询场景,涵盖土壤特征、建筑风格、车牌形状等数百种判别性属性。我们评估了LLM-as-a-judge和VLM-as-a-judge两种评分策略,发现Qwen 3作为裁判与人类专家评分相关性最高。结果表明,尽管闭源大模型如Gemini和GPT 5在定位上媲美人类,但在生成可审计的推理链上仍显著落后。开源小模型如Llama和Qwen在该任务中表现惨败,仅略优于一个已知位置却无视觉信息的基线幻觉推理链。我们认为这暴露了模型从高分辨率图像中提取细粒度视觉属性的能力局限。本研究公开全部数据供社区使用。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) are good at recognizing the global location of a photograph -- their geolocation prediction accuracy rivals the best human experts. But many VLMs are startlingly bad at \textit{explaining} which image evidence led to their prediction, even when their location prediction is correct. In this paper, we introduce GeoRC, the first benchmark for geolocation reasoning chains sourced directly from Champion-tier GeoGuessr experts, including the reigning world champion. This benchmark consists of 800 ``ground truth'' reasoning chains across 500 query scenes from GeoGuessr maps, with expert chains addressing hundreds of different discriminative attributes, such as soil properties, architecture, and license plate shapes. We evaluate LLM-as-a-judge and VLM-as-a-judge strategies for scoring VLM-generated reasoning chains against our expert reasoning chains and find that Qwen 3 LLM-as-a-judge correlates best with human-expert scoring. Our benchmark reveals that while large, closed-source VLMs such as Gemini and GPT 5 rival human experts at predicting locations, they still lag behind human experts when it comes to producing auditable reasoning chains. Small open-weight VLMs such as Llama and Qwen catastrophically fail on our benchmark -- they perform only slightly better than a baseline in which an LLM hallucinates a reasoning chain with oracle knowledge of the photo location but \textit{no visual information at all}. We believe the gap between human experts and VLMs on this task points to VLM limitations at extracting fine-grained visual attributes from high resolution images. We open source our benchmark for the community to use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。