让多模态大模型更准地从图表图像中提取数据
Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training Framework

- 从人类读图习惯出发,设计渐进式数据提取流程
- 在无标签图表上实现7B模型的顶尖数值准确率
- 适合需要高效可靠图表数据提取的研究者
图表数据提取能从图表图像中还原数据表,对可复现性、分析与重用至关重要。现有交互工具虽可靠但繁琐,混合式系统效率高却泛化能力差。近期多模态大语言模型(MLLMs)提供了统一接口,但其在无可见标签时精确提取数值的能力仍不明确。我们构建了一个包含多种真实世界无标签图表的基准测试集以评估该能力。结果表明,当前MLLMs虽能可靠重建表格结构,但在数值恢复上表现不佳。为此,我们从人机协同视角重新思考提取过程,提出类人渐进学习框架。该训练方法显著提升数值准确性,在7B参数模型上达到当前最优表现。用户研究进一步验证,该模型能有效支持混合式工作流,实现可靠的数据提取。
原文摘要 · Abstract (English)
Chart data extraction, which reverse-engineers data tables from chart images, is essential for reproducibility, analysis, retrieval, and redesign. Existing interactive tools are reliable but tedious, and mixed-initiative systems, while more efficient, lack generalizability. Recent multimodal large language models (MLLMs) offer a unified interface for chart interpretation, yet their ability to extract accurate data tables, especially without visible labels, remains unclear. We build a benchmark featuring diverse real-world charts without data labels to evaluate this capability. Results show that, while current MLLMs reliably reconstruct table structures, they struggle with precise value recovery. To address this, we revisit chart data extraction from a human-centered perspective and argue that extraction should follow a progressive learning process similar to how people read charts. Our training framework substantially improves numerical accuracy, achieving state-of-the-art performance with a 7B-parameter model. A user study further shows that our model effectively supports mixed-initiative workflows for reliable chart data extraction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。