评测大模型读图能力,覆盖三类地图和六类主题
MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering
- 构建包含14706组问答的MapIQ基准数据集
- 发现模型对地图设计变化敏感,依赖地理知识
- 适合研究地图理解、多模态推理的学者使用
近年来,多模态大语言模型(MLLMs)在读取数据可视化(如柱状图、散点图)方面取得进展。近期研究转向地图视觉问答(Map-VQA),但主要聚焦于专题地图(choropleth maps),涵盖的主题和分析任务有限。为填补这一空白,我们提出MapIQ,一个包含14,706组问题-答案对的基准数据集,覆盖三类地图类型:专题地图、变形地图(cartograms)和比例符号地图(proportional symbol maps),涉及住房、犯罪等六个不同主题。我们评估了多个MLLM在六种视觉分析任务上的表现,对比其性能与人类基准。此外,通过改变地图设计(如颜色方案、图例形式、移除地图元素)的实验,揭示了MLLM的鲁棒性、对设计变化的敏感度、对内部地理知识的依赖性,以及提升Map-VQA性能的潜在方向。
原文摘要 · Abstract (English)
Recent advancements in multimodal large language models (MLLMs) have driven researchers to explore how well these models read data visualizations, e.g., bar charts, scatter plots. More recently, attention has shifted to visual question answering with maps (Map-VQA). However, Map-VQA research has primarily focused on choropleth maps, which cover only a limited range of thematic categories and visual analytical tasks. To address these gaps, we introduce MapIQ, a benchmark dataset comprising 14,706 question-answer pairs across three map types: choropleth maps, cartograms, and proportional symbol maps spanning topics from six distinct themes (e.g., housing, crime). We evaluate multiple MLLMs using six visual analytical tasks, comparing their performance against one another and a human baseline. An additional experiment examining the impact of map design changes (e.g., altered color schemes, modified legend designs, and removal of map elements) provides insights into the robustness and sensitivity of MLLMs, their reliance on internal geographic knowledge, and potential avenues for improving Map-VQA performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。