用大模型自动提取混凝土材料数据,一小时处理近万条高质记录。
Large language model-enabled automated data extraction for concrete materials informatics

- 基于大语言模型构建自动化数据提取流水线,适配多种模型。
- 在27000+论文中提取近9000条数据,属性提取F1达0.98。
- 可扩展至其他材料领域,助力材料信息学数据基建。
数据驱动的材料发现受高质量、可访问实验数据稀缺的制约。本文提出一种通用的大语言模型(LLM)驱动的数据提取与结构化流程,以混凝土材料为例,解决非结构化文献中的数据提取难题。该流程在多种LLM上表现稳健,对成分-工艺-性能等属性的提取F1最高达0.98。在1小时内,从超过27,000篇文献中筛选并提取近9,000条高质量记录,每条含百余项属性,构建了目前最大的开放实验室混合水泥混凝土数据库。机器学习分析表明,大规模、多样化、信息丰富的数据集对提升模型在分布内准确率和未见材料的泛化能力至关重要。该方法可轻松拓展至其他材料领域,加速材料信息学可扩展数据基础设施的建设。
原文摘要 · Abstract (English)
The promise of data-driven materials discovery remains constrained by the scarcity of large, high-quality, and accessible experimental datasets. Here, we introduce a generalizable large language model (LLM)-powered pipeline for automated extraction and structuring of materials data from unstructured scientific literature, using concrete materials as a representative and particularly challenging example. The pipeline exhibits robust performance across a broad range of LLMs and achieves an $F_1$ score of up to 0.98 for diverse composition--process--property attributes. Within one hour, it extracts nearly 9,000 high-quality records with over 100 attributes from a corpus screened from more than 27,000 publications, enabling the construction of the largest open laboratory database for blended cement concrete. Machine learning analyses underscore the importance of large, diverse, and information-rich datasets for enhancing both in-distribution accuracy and out-of-distribution generalization to unseen materials. The proposed pipeline is readily adaptable to other materials domains and accelerates the development of scalable data infrastructures for materials informatics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。