构建图文问答基准,评估大模型理解信息图设计元素的能力。
InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts
- 对比同一数据的图文与普通图表,测试模型对设计元素的理解。
- 20个大模型在信息图上表现下降,尤其在隐喻类问题上差距显著。
- 适合研究多模态推理、信息可视化和模型可解释性的学者使用。
理解包含设计驱动视觉元素(如象形图、图标)的信息图需要结合视觉识别与推理能力,这对多模态大语言模型(MLLMs)构成挑战。然而,现有视觉问答基准因缺乏配对的普通图表与基于视觉元素的问题,难以有效评估该能力。为此,我们提出InfoChartQA,一个用于评估MLLMs在信息图理解上的基准。该数据集包含5,642对信息图与普通图表,共享相同底层数据但呈现方式不同。我们设计了基于视觉元素的问题,以捕捉其独特设计与传播意图。对20个MLLM的评估显示,模型在信息图上的表现显著下降,尤其在涉及隐喻的视觉元素问题上。成对的图表支持细粒度错误分析与消融实验,揭示了提升MLLM信息图理解能力的新方向。InfoChartQA已开源:https://github.com/CoolDawnAnt/InfoChartQA。
原文摘要 · Abstract (English)
Understanding infographic charts with design-driven visual elements (e.g., pictograms, icons) requires both visual recognition and reasoning, posing challenges for multimodal large language models (MLLMs). However, existing visual-question answering benchmarks fall short in evaluating these capabilities of MLLMs due to the lack of paired plain charts and visual-element-based questions. To bridge this gap, we introduce InfoChartQA, a benchmark for evaluating MLLMs on infographic chart understanding. It includes 5,642 pairs of infographic and plain charts, each sharing the same underlying data but differing in visual presentations. We further design visual-element-based questions to capture their unique visual designs and communicative intent. Evaluation of 20 MLLMs reveals a substantial performance decline on infographic charts, particularly for visual-element-based questions related to metaphors. The paired infographic and plain charts enable fine-grained error analysis and ablation studies, which highlight new opportunities for advancing MLLMs in infographic chart understanding. We release InfoChartQA at https://github.com/CoolDawnAnt/InfoChartQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。