用科学文献训练的元素嵌入模型,提升材料性能预测准确率
Semantic Embeddings of Chemical Elements for Enhanced Materials Inference and Discovery
- 基于129万篇合金论文训练专用BERT模型,提取元素语义特征
- 在钛合金等材料上预测精度最高提升23%,优于传统经验描述符
- 适合材料发现与优化研究者,尤其关注合金设计与性质预测
我们提出一种生成通用化学元素语义嵌入的框架,以推动材料推断与发现。该框架利用ElementBERT——一个在129万篇合金相关科学论文摘要上训练的领域专用BERT模型,捕捉合金特有的潜在知识和上下文关系。这些语义嵌入作为稳健的元素描述符,在多个下游任务中显著优于传统经验描述符,包括预测力学与相变性能、分类相结构,以及通过贝叶斯优化优化材料性能。在钛合金、高熵合金和形状记忆合金上的应用显示,预测准确率最高提升23%。结果表明,ElementBERT通过编码专业合金知识,超越通用BERT模型。该框架将科学文献中的上下文洞见与定量推断结合,加速先进材料的发现与优化,未来可拓展至其他材料类别。
原文摘要 · Abstract (English)
We present a framework for generating universal semantic embeddings of chemical elements to advance materials inference and discovery. This framework leverages ElementBERT, a domain-specific BERT-based natural language processing model trained on 1.29 million abstracts of alloy-related scientific papers, to capture latent knowledge and contextual relationships specific to alloys. These semantic embeddings serve as robust elemental descriptors, consistently outperforming traditional empirical descriptors with significant improvements across multiple downstream tasks. These include predicting mechanical and transformation properties, classifying phase structures, and optimizing materials properties via Bayesian optimization. Applications to titanium alloys, high-entropy alloys, and shape memory alloys demonstrate up to 23% gains in prediction accuracy. Our results show that ElementBERT surpasses general-purpose BERT variants by encoding specialized alloy knowledge. By bridging contextual insights from scientific literature with quantitative inference, our framework accelerates the discovery and optimization of advanced materials, with potential applications extending beyond alloys to other material classes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。