评测大模型生成图表注释的能力,发现具体指令能提升质量。
ChartAnno: Evaluating MLLMs for Chart Annotation Generation

- 构建1200张真实图表数据集,含代码与三类注释指令。
- 专有模型表现更优,具体指令使注释质量显著提升。
- 图像输入增益有限,设计类任务受益明显,适合研究图文理解者。
多模态大语言模型在图表理解、生成和编辑方面已取得显著进展,但其对现有图表进行注释的能力仍待深入探索。图表注释是一项常见且具有挑战性的沟通任务,要求模型推断意图、理解图表语义,并合理添加文本或图形元素。为此,我们提出ChartAnno,一个用于评估多模态大语言模型在图表注释生成任务上的基准测试集,包含1200张真实世界图表,每张图配有代码及三类不同详细程度的注释指令。我们评估了10个代表性多模态大模型,在两种主要输入设置下:(1)仅使用图表代码,(2)同时使用图表代码与图表图像,并额外进行了仅图像的消融实验。结果表明,专有模型整体表现更优,尽管大规模开源模型正在缩小差距;更具体的指令可显著提升注释质量,而推断抽象意图仍是当前模型的最大难点。提供图表图像带来的总体增益有限,改进主要体现在设计相关指标上。这些发现突显了图表注释生成作为一项需要语义扎根与有效注释设计的挑战性任务。代码与数据将在后续版本中发布。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have made significant progress in chart understanding, generation, and editing, but their ability to annotate existing charts remains underexplored. Annotating charts is a common yet challenging communicative task, requiring models to infer intended messages, interpret chart semantics, and place appropriate textual or graphical elements. To address this gap, we introduce ChartAnno, a benchmark for evaluating MLLMs on chart annotation generation. It contains 1,200 real-world charts with paired code and annotation instructions across three levels of instruction specificity. We evaluate 10 representative MLLMs under two primary input settings: (1) chart code alone and (2) both chart code and chart image, and further include a chart image-only ablation study. Results show that proprietary models remain stronger overall, although large-scale open-source models narrow the gap. More specific instructions improve annotation quality, while inferring abstract intent remains most difficult for current MLLMs. Providing chart images brings limited overall gains, with improvements mainly appearing in design-related metrics. These findings highlight chart annotation generation as a challenging task requiring semantic grounding and effective annotation design. Code and data will be released in a future version.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。