构建科学文献图表新基准,提升模型真实理解能力
Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature
- 聚焦真实科研论文中的复杂流程图,扩充图表类型与语境信息
- 涵盖37,607张高质量图表与5,629道精心设计的问答题
- 引入类人类考试评估框架,适合研究多模态理解的学者使用
科学文献中的图表常包含多图、流程图、结构图等复杂视觉元素。现有基准存在图表类型单一、问题模板化、评估方法不充分等问题,导致模型性能虚高。为此,我们提出新基准SCI-CQA,重点关注常被忽略的流程图。从近十年15个顶级计算机科学会议论文中收集202,760张图文对,经严格筛选后保留37,607张高质量图表及上下文信息。SCI-CQA包含5,629道精心设计的问题,涵盖客观与开放题。同时提出高效标注流程,显著降低人工成本。最后,通过上下文感知分析,揭示上下文在解答难题中的关键作用。
原文摘要 · Abstract (English)
Scientific Literature charts often contain complex visual elements, including multi-plot figures, flowcharts, structural diagrams and etc. Evaluating multimodal models using these authentic and intricate charts provides a more accurate assessment of their understanding abilities. However, existing benchmarks face limitations: a narrow range of chart types, overly simplistic template-based questions and visual elements, and inadequate evaluation methods. These shortcomings lead to inflated performance scores that fail to hold up when models encounter real-world scientific charts. To address these challenges, we introduce a new benchmark, Scientific Chart QA (SCI-CQA), which emphasizes flowcharts as a critical yet often overlooked category. To overcome the limitations of chart variety and simplistic visual elements, we curated a dataset of 202,760 image-text pairs from 15 top-tier computer science conferences papers over the past decade. After rigorous filtering, we refined this to 37,607 high-quality charts with contextual information. SCI-CQA also introduces a novel evaluation framework inspired by human exams, encompassing 5,629 carefully curated questions, both objective and open-ended. Additionally, we propose an efficient annotation pipeline that significantly reduces data annotation costs. Finally, we explore context-based chart understanding, highlighting the crucial role of contextual information in solving previously unanswerable questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。