测试大模型在地图推理上的表现,发现普遍不及人类。
MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models
- 构建跨文本、API和视觉的多任务地图评测基准
- 30个模型平均准确率不足67%,距人类差20%以上
- 适合研究地理空间智能与导航系统的学者参考
近期基础模型在自主工具使用和推理方面取得进展,但在基于地图的推理能力仍不充分。为此,我们提出MapEval,一个涵盖180个城市、54个国家的基准评测体系,包含700道多选题,覆盖空间关系、导航、行程规划和真实地图交互等任务。不同于以往仅针对简单位置查询的评测,MapEval要求模型处理长上下文推理、API调用及视觉地图分析,是目前最全面的地理空间智能评估框架。对30个基础模型(包括Claude-3.5-Sonnet、GPT-4o、Gemini-1.5-Pro)的评估显示,无一模型准确率超过67%,开源模型表现显著更差,所有模型均比人类表现落后20%以上。结果揭示了模型在距离、方向、路径规划及地点特定推理方面的明显短板,凸显提升地理空间智能的迫切需求。相关资源已公开:https://mapeval.github.io/。
原文摘要 · Abstract (English)
Recent advancements in foundation models have improved autonomous tool usage and reasoning, but their capabilities in map-based reasoning remain underexplored. To address this, we introduce MapEval, a benchmark designed to assess foundation models across three distinct tasks - textual, API-based, and visual reasoning - through 700 multiple-choice questions spanning 180 cities and 54 countries, covering spatial relationships, navigation, travel planning, and real-world map interactions. Unlike prior benchmarks that focus on simple location queries, MapEval requires models to handle long-context reasoning, API interactions, and visual map analysis, making it the most comprehensive evaluation framework for geospatial AI. On evaluation of 30 foundation models, including Claude-3.5-Sonnet, GPT-4o, and Gemini-1.5-Pro, none surpass 67% accuracy, with open-source models performing significantly worse and all models lagging over 20% behind human performance. These results expose critical gaps in spatial inference, as models struggle with distances, directions, route planning, and place-specific reasoning, highlighting the need for better geospatial AI to bridge the gap between foundation models and real-world navigation. All the resources are available at: https://mapeval.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。