arXiv:2605.06021cs.CVcs.DL2026-05

用AI批量从科学图表中提取数据,准确率远超传统方法。

PlotPick: AI-powered batch extraction of numerical data from scientific figures

论文配图:PlotPick: AI-powered batch extraction of numerical data from scientific figures
图 1 · 摘自论文原文
  • 利用视觉语言模型自动识别图表并转为表格数据
  • 在300张图上召回率达88%-96%,显著高于旧方法
  • 特别擅长处理训练数据外的箱线图等复杂图表

系统综述和元分析常需从图表中获取数值数据,但人工提取效率低且难以扩展。我们提出PlotPick,一个开源工具,使用视觉语言模型(VLMs)批量提取科学图表中的结构化表格数据。我们在两个主流图表转表格基准(ChartX和PlotQA)上评估了三家供应商提供的六种VLMs,并与专用图表转表格模型DePlot对比。所有六种VLMs在两个基准上均优于DePlot。在ChartX(限于柱状图、折线图、箱线图和直方图;n=300)上,VLMs召回率达88%-96%,而DePlot仅71%。在PlotQA(n=529)上,VLMs的RMSF1达86%-99%,优于DePlot的94%。差距在专用模型未覆盖的图表类型上最大:箱线图上,DePlot仅24% RMSF1,而VLMs达83%-97%。PlotPick可在https://plotpick.streamlit.app获取。

原文摘要 · Abstract (English)

Systematic reviews and meta-analyses frequently require numerical data that authors report only as figures, yet manual digitisation is slow and does not scale. We present PlotPick, an open-source tool that uses vision-language models (VLMs) to batch-extract structured tabular data from scientific figures. We evaluate six VLMs from three providers on two established chart-to-table benchmarks (ChartX and PlotQA) and compare against the dedicated chart-to-table model DePlot. All six VLMs outperform DePlot on both benchmarks. On ChartX (restricted to bar charts, line charts, box plots, and histograms; n=300), VLMs achieve 88-96% recall versus 71% for DePlot. On PlotQA (n=529), VLMs achieve 86-99% RMSF1 versus 94% for DePlot. The gap is largest on chart types absent from the dedicated models' training data: on box plots, DePlot achieves 24% RMSF1 while VLMs achieve 83-97%. PlotPick is available at https://plotpick.streamlit.app.

图表解析数据提取视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。