arXiv:2512.03558cs.CVcs.CL2025-12中稿 · SIGSPATIAL 2025被引 1

构建首个地图理解评测基准,检验大模型看地图能力

CartoMapQA: A Fundamental Benchmark Dataset Evaluating Vision-Language Models on Cartographic Map Understanding

  • 设计包含2000+样本的问答数据集,覆盖符号识别到路径推理
  • 发现模型在地理空间推理和地图语义理解上存在明显短板
  • 适合研究地图理解、导航系统或城市规划相关AI的开发者使用

视觉语言模型(LVLMs)为融合视觉与文本信息提供了新可能,但其对地图的解读能力仍待探索。本文提出CartoMapQA,一个专用于评估LVLM地图理解能力的基准数据集,包含超过2000个样本,每个样本包含一张地图、一个问题(开放或选择题)及真实答案。任务涵盖符号识别、信息提取、比例尺理解与路径推理等低、中、高阶地图理解技能。对开源与商用模型的评估显示,模型普遍存在地图语义理解困难、地理空间推理能力弱、易受光学字符识别(OCR)错误影响等问题。通过定位这些缺陷,CartoMapQA为改进LVLM架构提供重要参考,助力发展更可靠的导航、地理搜索与城市规划类应用。代码与数据已开源:https://github.com/ungquanghuy-kddi/CartoMapQA.git

原文摘要 · Abstract (English)

The rise of Visual-Language Models (LVLMs) has unlocked new possibilities for seamlessly integrating visual and textual information. However, their ability to interpret cartographic maps remains largely unexplored. In this paper, we introduce CartoMapQA, a benchmark specifically designed to evaluate LVLMs' understanding of cartographic maps through question-answering tasks. The dataset includes over 2000 samples, each composed of a cartographic map, a question (with open-ended or multiple-choice answers), and a ground-truth answer. These tasks span key low-, mid- and high-level map interpretation skills, including symbol recognition, embedded information extraction, scale interpretation, and route-based reasoning. Our evaluation of both open-source and proprietary LVLMs reveals persistent challenges: models frequently struggle with map-specific semantics, exhibit limited geospatial reasoning, and are prone to Optical Character Recognition (OCR)-related errors. By isolating these weaknesses, CartoMapQA offers a valuable tool for guiding future improvements in LVLM architectures. Ultimately, it supports the development of models better equipped for real-world applications that depend on robust and reliable map understanding, such as navigation, geographic search, and urban planning. Our source code and data are openly available to the research community at: https://github.com/ungquanghuy-kddi/CartoMapQA.git

地图理解视觉语言模型评测基准地理信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。