arXiv:2603.28902cs.AI2026-03中稿 · ACL被引 1

首个跨图表对比理解的大规模基准,推动多图分析研究

ChartDiff: A Large-Scale Benchmark for Comprehending Pairs of Charts

  • 构建8541对图表数据集,支持跨图趋势、波动与异常对比
  • 通用模型在人类评估中表现最佳,但指标得分与实际质量不匹配
  • 多系列图表仍难处理,端到端模型对绘图工具差异鲁棒

图表是分析推理的核心,但现有图表理解基准几乎仅关注单图解读,缺乏对多图对比推理的支持。为此,我们提出ChartDiff,首个面向跨图表对比总结的大规模基准。该数据集包含8,541对图表,覆盖多样数据源、图表类型与视觉风格,每对均配有大语言模型生成并经人工验证的摘要,描述趋势、波动与异常差异。我们基于ChartDiff评估通用型、专有型及流水线式模型。结果表明,前沿通用模型在基于GPT的评价中表现最优,而专有与流水线方法虽获更高ROUGE分数,却在人类评估中表现较差,揭示了词面重叠与实际摘要质量之间的明显偏差。进一步发现,多系列图表对各类模型仍是挑战,而强端到端模型对绘图库差异相对鲁棒。整体表明,当前视觉-语言模型在多图推理上仍面临显著挑战,并确立ChartDiff作为推进多图理解研究的新基准。

原文摘要 · Abstract (English)

Charts are central to analytical reasoning, yet existing benchmarks for chart understanding focus almost exclusively on single-chart interpretation rather than comparative reasoning across multiple charts. To address this gap, we introduce ChartDiff, the first large-scale benchmark for cross-chart comparative summarization. ChartDiff consists of 8,541 chart pairs spanning diverse data sources, chart types, and visual styles, each annotated with LLM-generated and human-verified summaries describing differences in trends, fluctuations, and anomalies. Using ChartDiff, we evaluate general-purpose, chart-specialized, and pipeline-based models. Our results show that frontier general-purpose models achieve the highest GPT-based quality, while specialized and pipeline-based methods obtain higher ROUGE scores but lower human-aligned evaluation, revealing a clear mismatch between lexical overlap and actual summary quality. We further find that multi-series charts remain challenging across model families, whereas strong end-to-end models are relatively robust to differences in plotting libraries. Overall, our findings demonstrate that comparative chart reasoning remains a significant challenge for current vision-language models and position ChartDiff as a new benchmark for advancing research on multi-chart understanding.

图表理解多图对比基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。