构建图表对齐评测基准,精准评估模型理解图表细节能力。
ChartAB: A Benchmark for Chart Grounding & Dense Alignment
- 设计专用JSON模板,量化表格提取与元素定位精度
- 引入双阶段推理流程,支持跨图表元素对齐与对比分析
- 揭示现有模型在图表理解中的感知偏差与幻觉问题
图表在可视化、推理、数据分析及思想交流中至关重要。然而,现有视觉语言模型(VLMs)仍难以准确感知细节,难以提取图表中的细粒度结构,导致其在多图表比较与推理方面能力受限。本文提出全新「ChartAlign基准(ChartAB)」,全面评估VLMs在图表接地任务中的表现,包括表格数据提取、可视化元素定位及属性识别,覆盖多种类型与复杂度的图表。我们设计了专用JSON模板,以精准计算各任务的评估指标。通过引入新颖的两阶段推理流程,该基准还可评估模型在跨图表间对齐与比较元素/属性的能力。对多个近期VLMs的评测分析揭示了其在图表理解中的感知偏见、弱点、鲁棒性不足及幻觉现象。这些发现凸显了当前模型在细粒度图表理解任务中的差异,并指明了亟需提升的具体能力。
原文摘要 · Abstract (English)
Charts play an important role in visualization, reasoning, data analysis, and the exchange of ideas among humans. However, existing vision-language models (VLMs) still lack accurate perception of details and struggle to extract fine-grained structures from charts. Such limitations in chart grounding also hinder their ability to compare multiple charts and reason over them. In this paper, we introduce a novel "ChartAlign Benchmark (ChartAB)" to provide a comprehensive evaluation of VLMs in chart grounding tasks, i.e., extracting tabular data, localizing visualization elements, and recognizing various attributes from charts of diverse types and complexities. We design a JSON template to facilitate the calculation of evaluation metrics specifically tailored for each grounding task. By incorporating a novel two-stage inference workflow, the benchmark can further evaluate VLMs capability to align and compare elements/attributes across two charts. Our analysis of evaluations on several recent VLMs reveals new insights into their perception biases, weaknesses, robustness, and hallucinations in chart understanding. These findings highlight the fine-grained discrepancies among VLMs in chart understanding tasks and point to specific skills that need to be strengthened in current models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。