arXiv:2510.05142cs.CLcond-mat.mtrl-sci2025-10被引 4

用多阶段大模型从文献中精准提取材料全链条信息,减少遗漏、保证可信。

Reliable End-to-End Material Information Extraction from the Literature with Source-Tracked Multi-Stage Large Language Models

  • 分阶段迭代提取+溯源追踪,提升信息准确性与可靠性。
  • 特征级与元组级F1均达0.96,微结构类别召回率提升超10%。
  • 适合构建高精度材料数据库,适用于各类材料体系研究者。

数据驱动的材料发现需要大规模实验数据集,但多数信息仍被困在非结构化文献中。现有提取方法通常仅关注有限特征,未涵盖成分-工艺-微观结构-性能的集成关系,难以支撑全面数据库建设。为此,我们提出基于大语言模型的多阶段信息提取流程,仅从实验报道材料文献中捕获47项特征,涵盖成分、工艺、微观结构和性能。该流程结合迭代提取与源追踪机制,显著提升准确率与可靠性。在特征级(独立属性)与元组级(互依赖特征)评估中,F1分数均约0.96。相比无源追踪的单次提取,本方法使微结构类别在特征级与元组级的F1分别提升10.0%与13.7%,并将100篇关于析出相多主元合金文献中漏掉的材料数从49降至13(漏率由12.4%降至3.3%)。该流程可实现高效、可扩展的文献挖掘,生成高精度、低遗漏、零误报的数据集,为机器学习与材料信息学提供可信输入。模块化设计可泛化至多种材料体系,支持全面材料信息提取。

原文摘要 · Abstract (English)

Data-driven materials discovery requires large-scale experimental datasets, yet most of the information remains trapped in unstructured literature. Existing extraction efforts often focus on a limited set of features and have not addressed the integrated composition-processing-microstructure-property relationships essential for understanding materials behavior, thereby posing challenges for building comprehensive databases. To address this gap, we propose a multi-stage information extraction pipeline powered by large language models, which captures 47 features spanning composition, processing, microstructure, and properties exclusively from experimentally reported materials. The pipeline integrates iterative extraction with source tracking to enhance both accuracy and reliability. Evaluations at the feature level (independent attributes) and tuple level (interdependent features) yielded F1 scores around 0.96. Compared with single-pass extraction without source tracking, our approach improved F1 scores of microstructure category by 10.0% (feature level) and 13.7% (tuple level), and reduced missed materials from 49 to 13 out of 396 materials in 100 articles on precipitate-containing multi-principal element alloys (miss rate reduced from 12.4% to 3.3%). The pipeline enables scalable and efficient literature mining, producing databases with high precision, minimal omissions, and zero false positives. These datasets provide trustworthy inputs for machine learning and materials informatics, while the modular design generalizes to diverse material classes, enabling comprehensive materials information extraction.

材料信息提取大模型文献挖掘数据可信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。