评测大模型生成图表描述的准确性和洞察力,发现现有模型普遍存在不足。
ChartFI: Benchmarking Faithfulness and Insightfulness of Chart Descriptions from Multimodal Large Language Models

- 基于四大维度构建高质量图文对数据集
- 896组复杂图表与语义丰富描述,覆盖真实场景
- 提出四类新指标,全面评估描述质量
图表描述对可访问性、跨模态检索及读者从复杂可视化中提取洞察至关重要。随着多模态大语言模型(MLLMs)被广泛用于自动化图表描述生成,一个关键问题浮现:这些模型生成的描述在多大程度上忠实且富有洞察?现有基准存在两大缺陷:数据集包含简单同质的图表和浅层的事实罗列式描述;现有评估指标无法捕捉描述质量的多维特性。为此,我们提出图表忠实性与洞察力基准(ChartFI-Bench)。首先总结高质量图表描述的四个核心维度:事实准确性、显著特征强调、领域知识引导、图表-文本互补性。基于此,构建包含896个图表-描述对的高质量基准,涵盖视觉复杂的图表与语义丰富的描述。同时设计四个对齐的评估指标:忠实性(Faithfulness)、覆盖率(Coverage)、信息量(Informativeness)和敏锐度(Acuity),系统评估描述在各维度的表现。主流MLLM实验验证了该框架的有效性,并揭示了现有模型的普遍弱点。
原文摘要 · Abstract (English)
Chart descriptions are essential for accessibility, cross-modal retrieval, and assisting readers in extracting insights from complex visualizations. As multimodal large language models (MLLMs) are increasingly adopted for automated chart description generation, a critical question arises: how faithfully and insightfully do these models actually describe charts? Current benchmarks fall short on two fronts: existing datasets consist of simple, homogeneous charts paired with shallow, fact-enumerating descriptions; and prevailing metrics fail to capture the multi-faceted nature of description quality. To address these gaps, we present the Chart Faithfulness and Insightfulness Benchmark (ChartFI-Bench). We first summarize four dimensions that characterize high-quality chart descriptions: factual accuracy, salient feature emphasis, domain-informed guidance, and chart-text complementarity. Guided by these dimensions, we construct a high-quality benchmark comprising 896 chart-description pairs, which feature visually complex charts and semantically rich descriptions. Furthermore, we design four aligned evaluation metrics -- Faithfulness, Coverage, Informativeness, and Acuity -- to systematically assess the quality of descriptions across these dimensions. Experiments conducted on mainstream MLLMs demonstrate the effectiveness of the proposed framework and reveal common weaknesses among existing models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。