arXiv:2409.00255cs.CVcs.AI2024-09

评测视觉语言模型对统计地图的问答能力,填补地图理解研究空白。

MAPWise: Evaluating Vision-Language Models for Advanced Map Queries

  • 构建跨三大地区的地图问答数据集,含1000道题与43种问题类型。
  • 发现现有模型在空间关系推理和颜色映射理解上存在明显不足。
  • 适合关注地图智能、多模态推理的研究者使用。

视觉语言模型(VLMs)在需要联合理解视觉与语言信息的任务中表现出色。一个极具潜力但尚未充分探索的应用是基于各类地图回答问题。本研究考察了VLMs在回答基于等级统计图(choropleth maps)问题方面的有效性,这类地图广泛用于数据分析与可视化。为推动该领域研究,我们引入了一个新的基于地图的问答基准,涵盖美国、印度、中国三个地理区域的地图,每张地图配1000个问题,共包含43种多样化的问题模板,要求模型具备对相对空间关系、复杂地图特征及多步推理的理解能力。数据集涵盖离散与连续值地图,包含不同颜色映射方式、类别排序与风格模式,支持全面分析。我们在该基准上评估多个VLMs的表现,揭示其能力短板,并为模型改进提供洞见。

原文摘要 · Abstract (English)

Vision-language models (VLMs) excel at tasks requiring joint understanding of visual and linguistic information. A particularly promising yet under-explored application for these models lies in answering questions based on various kinds of maps. This study investigates the efficacy of VLMs in answering questions based on choropleth maps, which are widely used for data analysis and representation. To facilitate and encourage research in this area, we introduce a novel map-based question-answering benchmark, consisting of maps from three geographical regions (United States, India, China), each containing 1000 questions. Our benchmark incorporates 43 diverse question templates, requiring nuanced understanding of relative spatial relationships, intricate map features, and complex reasoning. It also includes maps with discrete and continuous values, encompassing variations in color-mapping, category ordering, and stylistic patterns, enabling comprehensive analysis. We evaluate the performance of multiple VLMs on this benchmark, highlighting gaps in their abilities and providing insights for improving such models.

地图理解视觉语言模型多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。