用大模型生成合成数据,提升无机材料设计效率。
Language Models Enable Data-Augmented Synthesis Planning for Inorganic Materials
- 直接调用通用大模型预测反应物与温度,无需微调。
- 生成2.8万条合成配方,使模型温度预测误差降至73℃以下。
- 适合材料研发、自动化实验平台快速构建合成方案。
当前无机合成规划主要依赖启发式方法或基于有限数据集训练的机器学习模型,限制了其泛化能力。我们证明,无需任务特定微调的语言模型(如GPT-4.1、Gemini 2.0 Flash、Llama 4 Maverick)可有效回忆合成条件,在1,000个独立反应测试集中,前1位原料预测准确率达53.8%,前5位达66.1%;对烧结与煅烧温度的预测平均绝对误差低于126℃,与专用回归方法相当。通过集成多个语言模型,预测准确率进一步提升,单次推理成本降低最高达70%。随后,我们利用语言模型生成28,548条合成反应配方,并与文献挖掘样本结合,预训练一个基于Transformer的模型SyntMTE。在混合数据集上微调后,该模型将烧结温度预测的均方绝对误差降至73℃,煅烧温度降至98℃,相比仅使用实验数据的基线模型性能提升最高达8.7%。在锂镧锆氧固态电解质的案例研究中,SyntMTE成功复现了实验观察到的掺杂依赖烧结趋势。该混合工作流实现了可扩展、数据高效的无机合成规划。
原文摘要 · Abstract (English)
Inorganic synthesis planning currently relies primarily on heuristic approaches or machine-learning models trained on limited datasets, which constrains its generality. We demonstrate that language models, without task-specific fine-tuning, can recall synthesis conditions. Off-the-shelf models, such as GPT-4.1, Gemini 2.0 Flash and Llama 4 Maverick, achieve a Top-1 precursor-prediction accuracy of up to 53.8 % and a Top-5 performance of 66.1 % on a held-out set of 1,000 reactions. They also predict calcination and sintering temperatures with mean absolute errors below 126 °C, matching specialized regression methods. Ensembling these language models further enhances predictive accuracy and reduces inference cost per prediction by up to 70 %. We subsequently employ language models to generate 28,548 synthetic reaction recipes, which we combine with literature-mined examples to pretrain a transformer-based model, SyntMTE. After fine-tuning on the combined dataset, SyntMTE reduces mean-absolute error in sintering temperature prediction to 73 °C and in calcination temperature to 98 °C. This strategy improves models by up to 8.7 % compared with baselines trained exclusively on experimental data. Finally, in a case study on Li7La3Zr2O12 solid-state electrolytes, we demonstrate that SyntMTE reproduces the experimentally observed dopant-dependent sintering trends. Our hybrid workflow enables scalable, data-efficient inorganic synthesis planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。