arXiv:2503.14607cs.CV2025-03被引 32

测试大模型能否像人一样看懂地图并给出导航指令。

Can Large Vision Language Models Read Maps Like a Human?

  • 构建首个面向人类可读像素地图的导航数据集MapBench。
  • 超过1600个路径规划问题,涵盖100张复杂地图场景。
  • 暴露大模型在空间推理与结构化决策上的显著短板。

本文提出MapBench——首个专为人类可读、基于像素的地图户外导航设计的数据集,源自复杂的路径规划场景。MapBench包含来自100张多样化地图的超过1600个像素空间路径规划问题。在该数据集中,大型视觉语言模型(LVLMs)需根据地图图像和起点终点地标查询,生成语言导航指令。每张地图配备地图空间场景图(MSSG)作为索引结构,实现自然语言转换与结果评估。实验表明,MapBench显著挑战当前最先进的LVLMs,无论采用零样本提示或链式思维(CoT)增强推理框架,均暴露出其在空间推理与结构化决策上的关键局限。对开源与闭源LVLM的评估验证了该数据集的高难度。相关代码与数据集已公开于https://github.com/taco-group/MapBench。

原文摘要 · Abstract (English)

In this paper, we introduce MapBench-the first dataset specifically designed for human-readable, pixel-based map-based outdoor navigation, curated from complex path finding scenarios. MapBench comprises over 1600 pixel space map path finding problems from 100 diverse maps. In MapBench, LVLMs generate language-based navigation instructions given a map image and a query with beginning and end landmarks. For each map, MapBench provides Map Space Scene Graph (MSSG) as an indexing data structure to convert between natural language and evaluate LVLM-generated results. We demonstrate that MapBench significantly challenges state-of-the-art LVLMs both zero-shot prompting and a Chain-of-Thought (CoT) augmented reasoning framework that decomposes map navigation into sequential cognitive processes. Our evaluation of both open-source and closed-source LVLMs underscores the substantial difficulty posed by MapBench, revealing critical limitations in their spatial reasoning and structured decision-making capabilities. We release all the code and dataset in https://github.com/taco-group/MapBench.

地图理解视觉语言模型空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。