arXiv:2512.08016cs.CVcs.AI2025-12中稿 · ICLR

构建地图推理新基准,测试模型多步空间理解能力

FRIEDA: Benchmarking Multi-Step Cartographic Reasoning in Vision-Language Models

  • 设计跨地图多步推理任务,涵盖拓扑、度量和方向三类空间关系
  • 最强模型准确率仅38.2%,远低于人类84.87%表现
  • 适用于评估视觉语言模型在地理信息处理中的空间智能

地图推理是通过整合图例、比例尺、方向、地图文字和几何图形来理解地理关系的能力。尽管对灾害响应和城市规划等关键任务至关重要,该能力仍缺乏系统评估。现有视觉语言模型研究常将地图视为图表的特例,忽略其分层符号体系与跨图空间关系。为此,我们提出FRIEDA,一个面向复杂开放性地图推理的基准测试。数据源自不同领域和地区的实际地图文档。依据地理信息系统文献,涵盖拓扑(相邻、包含、相交、等于)、度量(距离)和方向(方位)三类空间关系。所有问题需多步推理,部分需跨图定位与关联。评估11个先进视觉语言模型,分别在直接设置(提供相关地图)与上下文设置(需先识别相关地图)下进行。最强模型Gemini-2.5-Pro和GPT-5-Think准确率分别为38.20%和37.20%,远低于人类84.87%。结果揭示了模型在多步地图推理上的显著差距,确立FRIEDA作为推动视觉语言模型空间智能发展的严格基准。

原文摘要 · Abstract (English)

Cartographic reasoning is the skill of interpreting geographic relationships by aligning legends, map scales, compass directions, map texts, and geometries across one or more map images. Although essential as a concrete cognitive capability and for critical tasks such as disaster response and urban planning, it remains largely unevaluated. Building on progress in chart and infographic understanding, recent large vision language model studies on map visual question-answering often treat maps as a special case of charts. In contrast, map VQA demands comprehension of layered symbology (e.g., symbols, geometries, and text labels) as well as spatial relations tied to orientation and distance that often span multiple maps and are not captured by chart-style evaluations. To address this gap, we introduce FRIEDA, a benchmark for testing complex open-ended cartographic reasoning in LVLMs. FRIEDA sources real map images from documents and reports in various domains and geographical areas. Following classifications in Geographic Information System (GIS) literature, FRIEDA targets all three categories of spatial relations: topological (border, equal, intersect, within), metric (distance), and directional (orientation). All questions require multi-step inference, and many require cross-map grounding and reasoning. We evaluate eleven state-of-the-art LVLMs under two settings: (1) the direct setting, where we provide the maps relevant to the question, and (2) the contextual setting, where the model may have to identify the maps relevant to the question before reasoning. Even the strongest models, Gemini-2.5-Pro and GPT-5-Think, achieve only 38.20% and 37.20% accuracy, respectively, far below human performance of 84.87%. These results reveal a persistent gap in multi-step cartographic reasoning, positioning FRIEDA as a rigorous benchmark to drive progress on spatial intelligence in LVLMs.

地图推理空间智能多步推理视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。