arXiv:2503.10857cs.GRcs.AI2025-03被引 6

用图形认知理论测试大模型看图表能力,发现其存在根本性短板。

Towards Understanding Graphical Perception in Large Multimodal Models

  • 基于人类图形认知理论构建自动化评测框架
  • 发现顶级模型在跨图表类型、视觉元素理解上严重不足
  • 适合关注多模态模型感知缺陷的研究者和开发者

尽管大型多模态模型(LMMs)在需要知识、推理与感知结合的复杂任务中表现优异,但我们意外发现它们在仅需感知能力的简单信息图任务上表现不佳。现有基准主要聚焦于综合能力的任务,难以揭示模型在感知方面的细微局限。为此,我们借鉴图形认知理论,提出一种评估框架,用于分析LMMs在图表中的感知能力差距。该框架通过自动化任务生成与响应评估设计,实现对多种图表类型、视觉元素和任务类型的全面、可控测试。我们使用该框架在三个粒度层次(图表、视觉元素、像素)上评估了当前最先进的LMMs。结果表明,包括GPT-4o在内的主流模型存在三大关键缺陷:(1)无法跨图表类型泛化;(2)无法理解基础视觉元素;(3)无法在图表内跨值交叉参考。这些发现为提升LMMs的感知能力提供了重要方向。评估框架与标注数据已公开于https://github.com/microsoft/lmm-graphical-perception。

原文摘要 · Abstract (English)

Despite the promising results of large multimodal models (LMMs) in complex vision-language tasks that require knowledge, reasoning, and perception abilities together, we surprisingly found that these models struggle with simple tasks on infographics that require perception only. As existing benchmarks primarily focus on end tasks that require various abilities, they provide limited, fine-grained insights into the limitations of the models' perception abilities. To address this gap, we leverage the theory of graphical perception, an approach used to study how humans decode visual information encoded on charts and graphs, to develop an evaluation framework for analyzing gaps in LMMs' perception abilities in charts. With automated task generation and response evaluation designs, our framework enables comprehensive and controlled testing of LMMs' graphical perception across diverse chart types, visual elements, and task types. We apply our framework to evaluate and diagnose the perception capabilities of state-of-the-art LMMs at three granularity levels (chart, visual element, and pixel). Our findings underscore several critical limitations of current state-of-the-art LMMs, including GPT-4o: their inability to (1) generalize across chart types, (2) understand fundamental visual elements, and (3) cross reference values within a chart. These insights provide guidance for future improvements in perception abilities of LMMs. The evaluation framework and labeled data are publicly available at https://github.com/microsoft/lmm-graphical-perception.

多模态模型图表理解感知能力评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。