arXiv:2510.01235cs.LGcond-mat.mtrl-sci2025-10被引 6

用AI自动从万篇论文中提取热电材料数据,构建最大规模可复现数据库。

Automated Extraction of Material Properties using LLM-based AI Agents

  • 设计多智能体LLM流水线,动态分配资源实现高效精准提取
  • 产出27822条带温变属性的热电数据,准确率高达F1=0.91
  • 提供可交互网页工具,适合材料研究者快速检索与分析

材料快速发现受限于缺乏将性能指标与结构背景关联的大规模机器可读数据集。现有数据库或规模小、人工标注,或偏向第一性原理结果,实验文献未被充分利用。本文提出一种基于大语言模型(LLM)的智能体驱动工作流,自主提取约1万篇全文科学论文中的热电与结构性质。该流程融合动态令牌分配、零样本多智能体提取及条件表格解析,在准确率与计算成本间取得平衡。在50篇精选论文上的基准测试显示,GPT-4.1在热电属性上达到最高准确率(F1=0.91),结构属性为0.82;而GPT-4.1 Mini仅需极低开销即达相近表现(F1=0.89和0.81),支持大规模部署。应用该流程,我们整理出27,822条温度依赖性属性记录,涵盖优值(ZT)、塞贝克系数、电导率、电阻率、功率因子和热导率,并附带晶系、空间群、掺杂策略等结构信息。数据分析重现了合金优于氧化物、p型掺杂更优等已知趋势,同时揭示更广泛的构效关系。为促进社区共享,我们发布交互式网页探索器,支持语义筛选、数值查询与CSV导出。本研究构建了迄今最大的LLM标注热电数据集,提供可复现且成本可控的提取流程,为超越热电领域的可扩展数据驱动材料发现奠定基础。

原文摘要 · Abstract (English)

The rapid discovery of materials is constrained by the lack of large, machine-readable datasets that couple performance metrics with structural context. Existing databases are either small, manually curated, or biased toward first principles results, leaving experimental literature underexploited. We present an agentic, large language model (LLM)-driven workflow that autonomously extracts thermoelectric and structural-properties from about 10,000 full-text scientific articles. The pipeline integrates dynamic token allocation, zeroshot multi-agent extraction, and conditional table parsing to balance accuracy against computational cost. Benchmarking on 50 curated papers shows that GPT-4.1 achieves the highest accuracy (F1 = 0.91 for thermoelectric properties and 0.82 for structural fields), while GPT-4.1 Mini delivers nearly comparable performance (F1 = 0.89 and 0.81) at a fraction of the cost, enabling practical large scale deployment. Applying this workflow, we curated 27,822 temperature resolved property records with normalized units, spanning figure of merit (ZT), Seebeck coefficient, conductivity, resistivity, power factor, and thermal conductivity, together with structural attributes such as crystal class, space group, and doping strategy. Dataset analysis reproduces known thermoelectric trends, such as the superior performance of alloys over oxides and the advantage of p-type doping, while also surfacing broader structure-property correlations. To facilitate community access, we release an interactive web explorer with semantic filters, numeric queries, and CSV export. This study delivers the largest LLM-curated thermoelectric dataset to date, provides a reproducible and cost-profiled extraction pipeline, and establishes a foundation for scalable, data-driven materials discovery beyond thermoelectrics.

材料发现LLM应用数据挖掘热电材料

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。