新基准ChartQAPro提升图表问答真实性和难度,挑战大模型理解能力。
ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering
- 构建157个真实数据源的多样化图表库,涵盖信息图与仪表板。
- 21个模型在新基准上平均性能下降34.7%,大模型表现显著下滑。
- 适合研究图表理解、视觉推理及多模态模型优化的研究者使用。
图表在数据分析中无处不在,但复杂分析需大量感知与认知努力。图表问答(CQA)系统通过让模型理解并推理图表数据来自动化这一过程。然而,现有基准如ChartQA缺乏现实多样性,且现代大视觉语言模型(LVLMs)在此上已出现性能饱和。为此,我们提出ChartQAPro,包含来自157个多样化来源的1,341张图表,覆盖多种图表类型(如信息图、仪表板),并设计1,948道问题,包括多项选择、对话式、假设性及无法回答的问题,更贴近真实场景。21个模型评估显示,LVLMs在ChartQAPro上性能大幅下降:例如Claude Sonnet 3.5在ChartQA上得分为90.5%,在ChartQAPro上仅55.81%。我们进一步通过错误分析与消融实验揭示关键挑战与改进方向。代码与数据已开源于https://github.com/vis-nlp/ChartQAPro。
原文摘要 · Abstract (English)
Charts are ubiquitous, as people often use them to analyze data, answer questions, and discover critical insights. However, performing complex analytical tasks with charts requires significant perceptual and cognitive effort. Chart Question Answering (CQA) systems automate this process by enabling models to interpret and reason with visual representations of data. However, existing benchmarks like ChartQA lack real-world diversity and have recently shown performance saturation with modern large vision-language models (LVLMs). To address these limitations, we introduce ChartQAPro, a new benchmark that includes 1,341 charts from 157 diverse sources, spanning various chart types, including infographics and dashboards, and featuring 1,948 questions in various types, such as multiple-choice, conversational, hypothetical, and unanswerable questions, to better reflect real-world challenges. Our evaluations with 21 models show a substantial performance drop for LVLMs on ChartQAPro; e.g., Claude Sonnet 3.5 scores 90.5% on ChartQA but only 55.81% on ChartQAPro, underscoring the complexity of chart reasoning. We complement our findings with detailed error analyses and ablation studies, identifying key challenges and opportunities for advancing LVLMs in chart understanding and reasoning. We release ChartQAPro at https://github.com/vis-nlp/ChartQAPro.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。