arXiv:2503.12326cs.CVcond-mat.mtrl-sci2025-03被引 9

用大模型自动提取论文图表数据,准确率超90%

Leveraging Vision Capabilities of Multimodal LLMs for Automated Data Extraction from Plots

  • 设计零样本提示链PlotExtract,无需微调即可解析图表
  • 对可提取图表的点定位精度超90%,误差低于5%
  • 适合需要批量处理科研图表数据的研究者

自动化从研究文本中提取数据进展迅速,大语言模型(LLMs)进一步推动了这一进程。然而,从论文图表中提取数据仍是一项复杂任务,长期依赖人工操作。我们发现,通过合理指令与工程化流程,当前多模态大模型具备准确提取图表数据的能力。该能力源于预训练模型本身,仅需采用一种名为PlotExtract的零样本链式提示策略即可实现,无需微调。本文展示了PlotExtract方法,并在合成及真实发表的图表上评估其性能。分析仅针对双轴图表。对于可提取的图表,PlotExtract在点定位上达到超过90%的精度(约90%召回率),x和y坐标误差均在5%或以下。结果表明,多模态大模型为高通量图表数据提取提供了可行路径,在多数场景下可替代现有手动方法。

原文摘要 · Abstract (English)

Automated data extraction from research texts has been steadily improving, with the emergence of large language models (LLMs) accelerating progress even further. Extracting data from plots in research papers, however, has been such a complex task that it has predominantly been confined to manual data extraction. We show that current multimodal large language models, with proper instructions and engineered workflows, are capable of accurately extracting data from plots. This capability is inherent to the pretrained models and can be achieved with a chain-of-thought sequence of zero-shot engineered prompts we call PlotExtract, without the need to fine-tune. We demonstrate PlotExtract here and assess its performance on synthetic and published plots. We consider only plots with two axes in this analysis. For plots identified as extractable, PlotExtract finds points with over 90% precision (and around 90% recall) and errors in x and y position of around 5% or lower. These results prove that multimodal LLMs are a viable path for high-throughput data extraction for plots and in many circumstances can replace the current manual methods of data extraction.

图表提取多模态模型数据自动化零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。