arXiv:2608.10677cs.CV2026-08中稿 · the 2nd Workshop o…

新基准Chartography评估专业图表理解能力,模型表现仅45%。

Chartography: A Benchmark for Professional Chart Understanding

论文配图:Chartography: A Benchmark for Professional Chart Understanding
图 1 · 摘自论文原文
  • 用真实职业场景图表和专家提问构建新评测集
  • 顶尖模型准确率仅45%,远低于现有基准
  • 适合研究图表视觉感知与领域知识的学者

医疗、工程、金融、制造和科学等领域专业人士常基于图表做出关键决策。现有图表评测集存在局限:以柱状图、折线图、饼图为主,推理链较短,且已接近饱和,前沿模型得分已达80-90%。我们提出Chartography,包含100个任务,图表源自真实职业实践,采用领域专用格式(标准评测集少见),问题由长期阅读图表的专业人士撰写,并经三位额外专家独立验证。对30种前沿模型配置的评估(每任务20次评分)显示,最佳配置平均通过率仅为45.0%,其余在9.0%-39.5%之间。失败主要集中在视觉感知:模型易忽略细微特征、误读稀疏标注坐标轴数值、处理投影三维几何错误,或违背图表中的领域惯例。所有任务、图像、溯源元数据及评估代码均已开源。

原文摘要 · Abstract (English)

Professionals across medicine, engineering, finance, manufacturing, and the sciences often make consequential decisions from charts. Existing chart benchmarks do not sufficiently measure this ability: they are dominated by bar, line, and pie formats, rely on shorter reasoning chains, and are nearing saturation, with frontier models already scoring 80-90%. We introduce Chartography, a benchmark of 100 tasks that pair charts drawn from professional practice, in domain-specific formats that standard chart benchmarks rarely include, with questions written by professionals who read these charts for a living and independently verified by three additional experts. In an evaluation of 30 frontier-model configurations (20 scored trials per task), the best configuration reaches only 45.0% mean pass@1; the remainder span 9.0-39.5%. Failures concentrate in visual perception: models can miss nuanced features, misread values along sparsely labeled axes, mishandle projected 3D geometry, and violate domain conventions encoded in the chart. We release all tasks, images, provenance metadata, and evaluation code.

图表理解评测基准视觉感知专业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。