构建真实地图的问答基准,评估模型空间推理能力。
MapVerse: A Benchmark for Geospatial Question Answering on Diverse Real-World Maps
- 基于1025张真实地图,人工标注1.18万条问答对。
- 现有模型在复杂空间推理任务上表现不足,差距明显。
- 适合研究地理认知、多模态推理与地图理解的学者使用。
地图是承载地理、人口、基础设施与环境模式等结构化知识的重要载体。对这些知识进行推理需要模型整合空间关系、视觉线索、现实背景和领域专长——当前大语言模型(LLMs)与视觉-语言模型(VLMs)仍难以稳定实现。然而,现有用于评估地图推理的基准数据集范围狭窄,局限于特定领域,且严重依赖人工生成内容(如大模型输出或流水线方法),难以反映真实的地理空间推理能力。为此,我们提出MapVerse,一个大规模真实地图基准。该数据集包含1,025张真实地图上的11,837条人类撰写的问题-答案对,覆盖10类地图类型及多种问题类别。它为地图阅读、解释与多模态推理提供了丰富评估场景。我们评估了10种前沿模型的表现,建立基线并量化推理差距。除整体性能外,还进行细粒度分类分析,考察模型在多维度上的推理表现,并探究影响推理结果的视觉因素。结果显示,尽管当前VLMs在分类类任务中表现良好,但开源与闭源模型在需复杂空间推理的高级任务上仍显著落后。
原文摘要 · Abstract (English)
Maps are powerful carriers of structured and contextual knowledge, encompassing geography, demographics, infrastructure, and environmental patterns. Reasoning over such knowledge requires models to integrate spatial relationships, visual cues, real-world context, and domain-specific expertise-capabilities that current large language models (LLMs) and vision-language models (VLMs) still struggle to exhibit consistently. Yet, datasets used to benchmark VLMs on map-based reasoning remain narrow in scope, restricted to specific domains, and heavily reliant on artificially generated content (outputs from LLMs or pipeline-based methods), offering limited depth for evaluating genuine geospatial reasoning. To address this gap, we present MapVerse, a large-scale benchmark built on real-world maps. It comprises 11,837 human-authored question-answer pairs across 1,025 maps, spanning ten diverse map categories and multiple question categories for each. The dataset provides a rich setting for evaluating map reading, interpretation, and multimodal reasoning. We evaluate ten state-of-the-art models against our benchmark to establish baselines and quantify reasoning gaps. Beyond overall performance, we conduct fine-grained categorical analyses to assess model inference across multiple dimensions and investigate the visual factors shaping reasoning outcomes. Our findings reveal that while current VLMs perform competitively on classification-style tasks, both open- and closed-source models fall short on advanced tasks requiring complex spatial reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。