arXiv:2510.26824cs.DLcs.AI2025-10被引 3

用AI从文献中自动提取材料合成方法,构建全球最大结构化数据库

LeMat-Synth: a multi-modal toolbox to curate broad synthesis procedure databases from scientific literature

  • 结合大模型与视觉模型,从文本和图表中自动解析合成流程
  • 建成包含5.8万条记录的数据库,覆盖35种方法、16类无机材料
  • 可应用于催化性能分析与超导临界温度验证,开源可用

材料科学中先进实验方法的广泛应用催生了海量的程序性知识,这些知识分散在数十年的科学文献中,以非结构化形式存在,难以系统分析。本文提出LeMat-Synth Parser,一个模块化、开源的多模态提取工具箱,利用大语言模型(LLMs)和视觉语言模型(VLMs),从出版物的文本和图表中自动提取合成方案与性能指标。将该工具应用于8.1万篇开放获取论文,构建了LeMat-Synth数据集,包含5.8万条合成程序,据我们所知是目前最大且最多样化的结构化无机材料合成数据集,涵盖35种合成方法和16类材料,基于领域特定本体。通过领域专家标注和可扩展的LLM-as-a-judge框架验证提取质量,并对多组模型进行基准测试以识别最优配置并刻画跨模型偏差。为展示其可扩展性,我们将其应用于两个不同领域:首先,关联合成方案与催化剂身份,分析氨分解文献中的热催化性能;其次,在1,384篇超导论文中交叉验证文本与图示报告的临界转变温度,再利用验证后的流程恢复样品系列中每种成分的临界转变温度。LeMat-Synth Parser与LeMat-Synth数据集已开源至GitHub与Hugging Face。

原文摘要 · Abstract (English)

Wide access to advanced experimental methods in materials science has given rise to an abundance of procedural knowledge, which is scattered across decades of scientific literature and recorded in unstructured formats that are challenging to analyze systematically. In this work, we present LeMat-Synth Parser, a modular, open-source, and multi-modal extraction toolbox that utilizes large language models (LLMs) and vision language models (VLMs) to automatically structure synthesis protocols and performance metrics extracted from both text and figures of publications. Applying LeMat-Synth Parser to 81K open-access publications, we curate LeMat-Synth, an extensive dataset of 58K synthesis procedures and to our knowledge the largest and most diverse structured inorganic materials synthesis dataset to date, covering 35 synthesis methods and 16 material classes based on a domain-specific ontology. We validate extraction quality against annotations by domain experts and a scalable LLM-as-a-judge framework, and benchmark a suite of models to identify optimal configurations and characterize cross-model biases. To demonstrate the extensibility of LeMat-Synth Parser, we apply it to two distinct domains. First, we link synthesis protocols and catalyst identity to thermocatalytic performance across a corpus of ammonia-decomposition publications. Second, we cross-validate text- and figure-reported critical transition temperatures across 1,384 superconductivity papers, then use the validated pipeline to recover the critical transition temperature for every composition in a sample series. We release LeMat-Synth Parser and the LeMat-Synth dataset openly on GitHub and Hugging Face

材料科学多模态知识提取数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。