用大模型自动提取文献中的金属有机框架合成条件,准确率超99%。
Reshaping MOFs text mining with a dynamic multi-agents framework of large language model
- 基于大模型动态解析论文和晶体代码,统一命名并生成结构化表格。
- 合成信息提取准确率达99%,缩写识别解决率94.1%,处理效率每篇仅9.6秒。
- 适合材料化学、自动化实验设计及文献数据挖掘研究者使用。
准确识别金属有机框架(MOFs)的合成条件对指导实验设计至关重要,但文献中相关信息常分散、不一致且难以解读。我们提出MOFh6,一个由大语言模型驱动的系统,可读取原始论文或晶体代码,将其转化为标准化的合成表格。该系统能跨段落关联相关描述,统一配体缩写与全称,并输出可直接使用的结构化参数。MOFh6实现99%的信息提取准确率,在五家主要出版商中解决94.1%的缩写问题,精度保持在0.93 ± 0.01。单篇全文处理耗时9.6秒,定位合成信息需36秒,处理100篇仅需4.24美元。通过将静态数据库查询替换为实时提取,MOFh6重塑了MOF合成研究范式,加速文献知识向实际合成协议的转化,推动可扩展的数据驱动材料发现。
原文摘要 · Abstract (English)
Accurately identifying the synthesis conditions of metal-organic frameworks (MOFs) is essential for guiding experimental design, yet remains challenging because relevant information in the literature is often scattered, inconsistent, and difficult to interpret. We present MOFh6, a large language model driven system that reads raw articles or crystal codes and converts them into standardized synthesis tables. It links related descriptions across paragraphs, unifies ligand abbreviations with full names, and outputs structured parameters ready for use. MOFh6 achieved 99% extraction accuracy, resolved 94.1% of abbreviation cases across five major publishers, and maintained a precision of 0.93 +/- 0.01. Processing a full text takes 9.6 s, locating synthesis descriptions 36 s, with 100 papers processed for USD 4.24. By replacing static database lookups with real-time extraction, MOFh6 reshapes MOF synthesis research, accelerating the conversion of literature knowledge into practical synthesis protocols and enabling scalable, data-driven materials discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。