构建跨语言多场景的图表解析基准,统一评估模型在真实图像中的表现。
ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats

- 设计多阶段人工-智能协作标注流程,确保数据可靠性。
- 覆盖8类图表,涵盖数字图与流程图等结构,评估3种视觉场景。
- 提出格式无关评估协议,支持不同输出结果的公平比较。
图表是传递定量与关系信息的主要媒介,但系统性评估图表解析模型仍具挑战。现有基准局限于少数图表类型,忽视流程图、思维导图等图示结构,且模型输出格式不一,数据集也缺乏实际中常见的印刷或手绘图像。为此,我们提出ChartArena,一个全面的双语基准,涵盖八类图表家族,包括数值图表与图示结构,每类在三种视觉场景下评估:数字渲染、打印照片和手绘照片。数据通过人机协同标注管道构建,并经多阶段人工验证以保证可靠性。为实现公平跨模型比较,我们设计了格式无关的评估协议,将异构输出映射到两个标准语义空间——归一化三元组视图与有向图视图,并使用结构感知指标评分。对26个主流多模态大模型的广泛评估显示三个一致发现:(i) 领先专有模型如Gemini 3.1 Pro整体领先,但最强开源系统正快速缩小差距;(ii) 文档解析模型对数值图表处理尚可,但在图示结构上表现显著落后;(iii) 专家图表解析器仍局限于狭窄图表类别。所有模型在雷达图和手绘场景中均面临较大挑战。这些发现表明ChartArena揭示了清晰的能力缺口,并为未来研究提供统一基础。ChartArena已公开于https://github.com/pspdada/ChartArena。
原文摘要 · Abstract (English)
Charts are a primary medium for conveying quantitative and relational information, yet systematically evaluating chart parsing models remains difficult. Existing benchmarks focus on narrow chart types and leave diagrammatic structures such as flowcharts and mind maps largely unaddressed, while models produce outputs in incompatible formats, and datasets rarely include the printed or hand-drawn images encountered in practice. To address these issues, we introduce ChartArena, a comprehensive bilingual benchmark covering eight chart families spanning both numeric charts and diagrammatic structures, each evaluated across three visual scenarios: digital renderings, printed photos, and hand-drawn photos. The dataset is built via a human-agent collaborative annotation pipeline with multi-stage human verification to ensure annotation reliability. To enable fair cross-model comparison, we further design a format-agnostic evaluation protocol that maps heterogeneous outputs into two canonical semantic spaces, a normalized triple view and a directed graph view, and scores them with structure-aware metrics. Through extensive evaluation of 26 leading MLLMs, we observe three consistent findings: (i) frontier proprietary models such as Gemini 3.1 Pro lead overall, yet the strongest open-source systems are rapidly closing the gap; (ii) document parsing models handle numeric charts reasonably but fall sharply behind on diagrammatic structures; and (iii) expert chart parsers remain limited to narrow chart families. Across all models, radar charts and hand-drawn scenarios stay especially challenging. These findings show that ChartArena exposes clear capability gaps and provides a unified foundation for future progress. ChartArena is publicly available at https://github.com/pspdada/ChartArena.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。