用多智能体自动提取文献中的复杂材料数据,构建可训练数据库
ComProScanner: A multi-agent based framework for composition-property structured data extraction from scientific literature
- 设计多智能体系统协同完成数据提取、验证与分类
- 在100篇文献上实现82%准确率,最优模型为DeepSeek-V3-0324
- 适合材料科学家快速构建机器学习训练数据集
自预训练大语言模型出现以来,从科学文本中提取结构化知识的方式相比传统方法发生革命性变化。然而,能够让用户构建、验证并可视化从文献中提取数据集的可用自动化工具仍十分稀缺。为此,我们开发了ComProScanner——一个自主的多智能体平台,用于提取、验证、分类和可视化可机读的化学成分与性质数据,并整合合成数据以创建全面数据库。我们在100篇期刊文章上评估该框架,对比了10种不同LLM(包括开源与专有模型),针对陶瓷压电材料的复杂组成及其对应的压电应变系数(d33)进行提取。DeepSeek-V3-0324表现最优,整体准确率达0.82。该框架提供了一个简单、用户友好且即用型工具包,可从文献中挖掘复杂实验数据,用于构建机器学习或深度学习数据集。
原文摘要 · Abstract (English)
Since the advent of various pre-trained large language models, extracting structured knowledge from scientific text has experienced a revolutionary change compared with traditional machine learning or natural language processing techniques. Despite these advances, accessible automated tools that allow users to construct, validate, and visualise datasets from scientific literature extraction remain scarce. We therefore developed ComProScanner, an autonomous multi-agent platform that facilitates the extraction, validation, classification, and visualisation of machine-readable chemical compositions and properties, integrated with synthesis data from journal articles for comprehensive database creation. We evaluated our framework using 100 journal articles against 10 different LLMs, including both open-source and proprietary models, to extract highly complex compositions associated with ceramic piezoelectric materials and corresponding piezoelectric strain coefficients (d33), motivated by the lack of a large dataset for such materials. DeepSeek-V3-0324 outperformed all models with a significant overall accuracy of 0.82. This framework provides a simple, user-friendly, readily-usable package for extracting highly complex experimental data buried in the literature to build machine learning or deep learning datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。