arXiv:2604.21344cs.CLcs.AI2026-04ACL

构建多图表问答数据集,评估模型跨图理解能力

Beyond Single Plots: A Benchmark for Question Answering on Multi-Charts

论文配图:Beyond Single Plots: A Benchmark for Question Answering on Multi-Charts
图 1 · 摘自论文原文
  • 构建包含2297个子图表的多图表问答数据集
  • 模型在人工问题上准确率比自动生成问题低27.4%
  • 提出新提示方法使准确率提升5.39%,适合图表分析研究者

图表广泛用于呈现复杂信息。在真实场景中获取有意义洞察,往往需要联合解读多个相关图表。目前对多图表图像理解的研究尚不充分。我们提出PolyChartQA,一个中等规模的多图表图像问答数据集。该数据集包含534张多图表图像(共2,297个子图表),源自同行评审的计算机科学论文,以及2,694个问答对。我们在PolyChartQA上评估了九种前沿多模态语言模型(MLMs)在不同问题类型、难度、来源及多图表结构特征下的表现。结果显示,模型在人工生成的问题上相比自动生成问题,准确率下降27.4%;采用我们提出的提示方法后,准确率提升5.39%。

原文摘要 · Abstract (English)

Charts are widely used to present complex information. Deriving meaningful insights in real-world contexts often requires interpreting multiple related charts together. Research on understanding multi-chart images has not been extensively explored. We introduce PolyChartQA, a mid-scale dataset specifically designed for question answering over multi-chart images. PolyChartQA comprises 534 multi-chart images (with a total of 2,297 sub-charts) sourced from peer-reviewed computer science research publications and 2,694 QA pairs. We evaluate the performance of nine state-of-the-art Multimodal Language Models (MLMs) on PolyChartQA across question type, difficulty, question source, and key structural characteristics of multi-charts. Our results show a 27.4% LLM-based accuracy (L-Accuracy) drop on human-authored questions compared to MLM-generated questions, and a 5.39% L-accuracy gain with our proposed prompting method.

多图表理解问答系统多模态模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。