arXiv:2607.09068cs.CVcs.AI2026-07被引 1

构建地图文档视觉推理基准,挑战模型真视觉理解能力

OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents

论文配图:OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents
图 1 · 摘自论文原文
  • 设计九类地图数据集,强制模型依赖视觉而非文本
  • 25个主流多模态模型最高仅75.03%准确率,暴露视觉理解短板
  • 提出视觉依赖指数(VDI),量化模型对图像的不可替代性

近期大视觉语言模型(LVLMs)的发展亟需可靠的复杂视觉推理基准。现有文档理解基准普遍存在的问题是:视觉内容可被文字替代,导致高分无需真正视觉感知。为此,本文提出OmniMapBench,旨在推动地图文档的视觉中心型推理研究。该基准包含1,603张地图文档,覆盖九个类别,共2,096组人工标注的问答对,系统评估从感知到多步视觉推理的多层次能力。为量化基准特性,提出简单有效的基准级指标——视觉依赖指数(VDI),即用与问题无关的描述替换图像后准确率的下降程度。OmniMapBench的VDI高于现有基准,验证其对不可约视觉推理的聚焦。对25个领先LVLM进行全面评估,表现差距显著,最优模型仅达75.03%准确率,凸显当前模型在复杂地图理解上的局限。本工作旨在推动LVLM在文档理解中的视觉推理进步。数据集与代码已公开于https://github.com/SIGMME/OmniMapBench。

原文摘要 · Abstract (English)

Recent advancements in LVLMs necessitate robust benchmarks for complex, visually grounded reasoning. A critical limitation is identified in many document understanding benchmarks: visual content is often reducible to text, enabling high performance without genuine visual grounding. To address this limitation, OmniMapBench is introduced to foster visual-centric reasoning for map documents. The benchmark comprises 2,096 manually annotated question-answer pairs across 1,603 map documents from nine categories. It is designed to probe a hierarchy of skills, ranging from perception to multi-step visual reasoning. To quantify benchmark properties, a simple yet effective benchmark-level metric is proposed: the Visual Dependency Index (VDI), defined as the accuracy drop when images are replaced with question-agnostic descriptions. OmniMapBench exhibits higher VDI than established benchmarks, which quantitatively validates its focus on irreducible visual reasoning. Comprehensive evaluations of 25 leading LVLMs are conducted on OmniMapBench. A significant performance gap is observed, with the top-performing model achieving only 75.03\% accuracy. This result underscores the challenges posed by OmniMapBench to current LVLMs. This work aims to catalyze progress in visual-centric reasoning for document understanding of LVLMs. The dataset and code are publicly available at https://github.com/SIGMME/OmniMapBench.

视觉推理地图理解多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。