arXiv:2601.12585cs.HCcs.AI2026-01被引 1

首次系统分析AI看图能力短板,揭示模型在复杂图表上的理解瓶颈。

Do MLLMs See What We See? Analyzing Visualization Literacy Barriers in AI Systems

  • 基于人类认知框架,编码4个顶尖模型309个错误回答,构建故障分类体系。
  • 模型在简单图表上表现良好,但在颜色密集、分段复杂的图表中准确率显著下降。
  • 发现两类机器特有障碍,为未来可视化AI设计提供关键改进方向。

多模态大语言模型(MLLMs)被越来越多用于解析可视化图表,但其失败原因尚不明确。本文首次系统分析了MLLM在可视化理解中的障碍,采用重构的可视化素养评估测试(reVLAT)基准与合成数据,对四个先进模型的309个错误响应进行以障碍为中心的开码分析,借鉴人类可视化素养研究策略。分析结果形成一个MLLM失败的分类体系,揭示了两类扩展既有以人类参与为基础框架的机器特有障碍。结果显示,模型在简单图表上表现良好,但在颜色密集、分段结构复杂的可视化中,常无法建立一致的比较推理。研究为未来可靠AI可视化助手的评估与设计提供了重要依据。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) are increasingly used to interpret visualizations, yet little is known about why they fail. We present the first systematic analysis of barriers to visualization literacy in MLLMs. Using the regenerated Visualization Literacy Assessment Test (reVLAT) benchmark with synthetic data, we open-coded 309 erroneous responses from four state-of-the-art models with a barrier-centric strategy adapted from human visualization literacy research. Our analysis yields a taxonomy of MLLM failures, revealing two machine-specific barriers that extend prior human-participation frameworks. Results show that models perform well on simple charts but struggle with color-intensive, segment-based visualizations, often failing to form consistent comparative reasoning. Our findings inform future evaluation and design of reliable AI-driven visualization assistants.

多模态模型可视化理解认知障碍

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。