arXiv:2506.00785cs.AIcs.CV2025-06EMNLP被引 24

构建多模态地理推理评测基准,揭示大模型在空间理解上的短板

GeoChain: Multimodal Chain-of-Thought for Geographic Reasoning

  • 基于146万张街景图设计21步推理链,覆盖视觉、空间等四类地理推理
  • 测试显示主流模型在复杂推理中定位准确率低,且推理过程不稳定
  • 适合研究多模态大模型地理认知能力的学者和开发者参考

本文提出GeoChain,一个用于评估多模态大语言模型(MLLMs)逐步地理推理能力的大规模基准。该基准利用146万张Mapillary街景图像,为每张图配以21步链式问题序列(共3000万以上问答对),引导模型从粗粒度属性推导到细粒度定位,涵盖视觉、空间、文化及精确地理定位四类推理任务,并标注难度。图像还附加150类语义分割标签和视觉可定位性评分。在2088张图像子集上对当前主流MLLMs(GPT-4.1变体、Claude 3.7、Gemini 2.5变体)进行测试发现,模型普遍存在视觉对齐偏差、推理波动大、复杂任务定位不准等问题。GeoChain提供了一套可靠的诊断方法,对推动MLLM在复杂地理推理中的进展至关重要。

原文摘要 · Abstract (English)

This paper introduces GeoChain, a large-scale benchmark for evaluating step-by-step geographic reasoning in multimodal large language models (MLLMs). Leveraging 1.46 million Mapillary street-level images, GeoChain pairs each image with a 21-step chain-of-thought (CoT) question sequence (over 30 million Q&A pairs). These sequences guide models from coarse attributes to fine-grained localization across four reasoning categories - visual, spatial, cultural, and precise geolocation - annotated by difficulty. Images are also enriched with semantic segmentation (150 classes) and a visual locatability score. Our benchmarking of contemporary MLLMs (GPT-4.1 variants, Claude 3.7, Gemini 2.5 variants) on a diverse 2,088-image subset reveals consistent challenges: models frequently exhibit weaknesses in visual grounding, display erratic reasoning, and struggle to achieve accurate localization, especially as the reasoning complexity escalates. GeoChain offers a robust diagnostic methodology, critical for fostering significant advancements in complex geographic reasoning within MLLMs.

地理推理多模态大模型评测链式思维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。