arXiv:2606.00065cs.IRcond-mat.mtrl-sci2026-06

用视觉语言模型从科学图表中精准提取材料数据,打通文献自动化分析最后一环。

Beyond Text and Tables: Vision-Language Model Integration in ComProScanner for Extracting Materials Data from Scientific Figures with High Accuracy

论文配图:Beyond Text and Tables: Vision-Language Model Integration in ComProScanner for Extracting Materials Data from Scientific Figures with High Accuracy
图 1 · 摘自论文原文
  • 引入视觉语言模型,从图表中自动识别材料成分与性能数据
  • Gemini-3-Flash-Preview在准确率和成本上表现最优,组合精度达0.97
  • 支持图表范围误差评估,更符合实际材料科学的物理意义

基于大语言模型的自动化材料数据提取已取得显著进展,但现有框架仍局限于文本与表格,忽略了大量仅以科学图表形式呈现的定量数据。本文扩展了端到端多智能体框架ComProScanner,新增原生视觉语言模型(VLM)支持的图表数据提取能力。引入FigureExtractor工具实现跨出版商的基于标题与关键词的图表筛选,并通过GraphExtractorTool将图表传入可配置的VLM,恢复其中的成分-性能配对信息。选取4种在LMArena Diagram排行榜上表现优异且每百万词输入成本低于1.50美元的VLM进行评估。在50篇压电陶瓷论文组成的$ d_{33} $测试语料库上,Gemini-3-Flash-Preview表现最佳,成分准确率为0.97,归一化F1得分为0.97,同时为四者中最经济高效。此外,提出基于范围的数值误差阈值,使对图表中数值属性的评估更具物理合理性。该工作使集成VLM的ComProScanner成为首个能统一处理文本、表格与图表的材料专用全自动文献挖掘平台。

原文摘要 · Abstract (English)

Automated extraction of materials composition-property data from scientific literature has advanced considerably with the development of large language model-based pipelines; however, existing frameworks remain limited to textual and tabular content, overlooking the substantial proportion of quantitative property data reported exclusively in scientific figures. Here, we extend ComProScanner, a fully end-to-end multi-agent framework for automated composition-property database construction, with a native vision-language model (VLM) based figure extraction capability. The extension introduces a FigureExtractor utility for caption-keyword-based figure filtering across all supported publishers, and a GraphExtractorTool agent that passes extracted figures to a configurable VLM to recover composition-property pairs from scientific charts and plots. Four VLMs are selected for evaluation on the basis of the LMArena Diagram leaderboard with an input cost criterion of less than \$1.50 per million tokens. Benchmarking on 50 piezoelectric ceramic articles from the established $d_{33}$ test corpus demonstrates that Gemini-3-Flash-Preview achieves the highest performance with a composition accuracy of 0.97 and a normalised F1 score of 0.97, whilst remaining the most cost-effective model among the four evaluated. We additionally introduce a range-based value error threshold parameter into the evaluation framework, providing a more physically meaningful assessment of numeric property values extracted from figures than exact value matching. These contributions establish VLM-integrated ComProScanner as the first materials-specific, fully automated, multimodal literature mining platform capable of extracting structured composition-property data from text, tables, and figures within a single unified pipeline.

材料数据视觉语言模型图表提取自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。