构建首个覆盖30类图表的多任务基准,评估模型对图表结构与数据的精准理解。
ChartAnchor: Chart Grounding with Structural-Semantic Fidelity
- 提出双向对齐的图表-代码生成与受控表格重建任务,兼顾视觉与数值一致性。
- 涵盖8000+图表-表格-代码三元组,覆盖真实世界与增强数据源中的30种图表类型。
- 适合关注科学、金融等领域的多模态模型结构化推理能力研究者使用。
近年来多模态大语言模型的发展凸显了对结构化图表理解能力评估基准的需求。图表定位指图表视觉外观与其结构语义之间的双向对齐,要求模型生成能忠实反映图表视觉与结构意图的符号化描述,并精确还原底层表格数据及其数值关系。该任务直接体现模型在数值推理、多模态对齐与结构重建方面的能力,具有重要实际应用价值。现有基准受限于图表类型单一、任务孤立及评估框架不完整,难以全面评估定位能力。为此,我们提出ChartAnchor,一个包含8000+图表-表格-代码三元组的综合性基准,覆盖30种来自真实世界与增强数据源的图表类型。ChartAnchor引入两个互补任务:图表到代码生成与受控图表到表格重建,实现视觉与数值保真度的交叉验证。采用多层次评估框架,融合语义验证、风格分析与感知指标,评估结构与内容层面的正确性。对MLLM的大规模实验揭示其在数值精度与代码合成上的关键缺陷,强调超越表层感知的结构性推理必要性。通过统一符号化与数据驱动的定位范式,ChartAnchor为图表定位建立了严格评估基础,为提升科学、金融与工业领域多模态模型能力提供实质性洞见。
原文摘要 · Abstract (English)
Recent advances in multimodal large language models (MLLMs) highlight the need for benchmarks that rigorously evaluate structured chart comprehension. Chart grounding refers to the bidirectional alignment between a chart's visual appearance and its structured semantics. This task requires models to produce a symbolic specification that faithfully captures the chart's visual and structural intent, while also recovering the underlying tabular data with precise values and relationships. Chart grounding directly reflects a model's capabilities in numerical reasoning, multimodal alignment, and structural reconstruction, and has several important real-world applications. Existing benchmarks, constrained by narrow chart diversity, isolated tasks, and incomplete evaluation frameworks, fail to holistically assess grounding. To address this, we propose ChartAnchor, a comprehensive benchmark of 8k+ chart-table-code triples spanning 30 chart types drawn from diverse real-world and augmented sources. ChartAnchor introduces two complementary tasks: chart-to-code generation and controlled chart-to-table reconstruction, enabling cross-validation of visual and numerical fidelity. A multi-level evaluation framework integrates semantic validation, stylistic analysis, and perceptual metrics to assess both structural and content-level correctness. Extensive experiments on MLLMs reveal critical limitations in numerical precision and code synthesis, emphasizing the need for structured reasoning beyond surface-level perception. By unifying symbolic and data-driven grounding, ChartAnchor establishes a rigorous foundation for chart grounding, offering meaningful insights for advancing MLLMs in scientific, financial, and industrial domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。