融合文本与物理特征,提升复杂材料性能预测精度
Information fusion strategy integrating pre-trained language model and contrastive learning for materials knowledge mining
- 结合材料文献文本与物理参数,用预训练模型和对比学习提取隐含知识
- 在钛合金和难熔多主元合金上分别达到R2=0.849和0.680
- 适合材料信息学、知识引导设计的科研人员参考
机器学习已推动材料设计革新,但受加工条件与微观结构影响,合金延展性等复杂性能预测仍具挑战,传统还原论方法难以量化。本文提出一种新型信息融合架构,将材料科学领域文本与定量物理描述符结合。框架采用MatSciBERT进行文本理解,并引入对比学习自动提取加工参数与微观结构特征的隐含知识。通过严格的消融实验与对比测试,模型表现优异,在钛合金验证集上R²达0.849,在难熔多主元合金测试集上R²达0.680。该系统性方法为定量描述符不全的复杂材料体系提供了全面的性能预测框架,奠定了知识引导材料设计与数据驱动材料发现的基础。
原文摘要 · Abstract (English)
Machine learning has revolutionized materials design, yet predicting complex properties like alloy ductility remains challenging due to the influence of processing conditions and microstructural features that resist quantification through traditional reductionist approaches. Here, we present an innovative information fusion architecture that integrates domain-specific texts from materials science literature with quantitative physical descriptors to overcome these limitations. Our framework employs MatSciBERT for advanced textual comprehension and incorporates contrastive learning to automatically extract implicit knowledge regarding processing parameters and microstructural characteristics. Through rigorous ablation studies and comparative experiments, the model demonstrates superior performance, achieving coefficient of determination (R2) values of 0.849 and 0.680 on titanium alloy validation set and refractory multi-principal-element alloy test set. This systematic approach provides a holistic framework for property prediction in complex material systems where quantitative descriptors are incomplete and establishes a foundation for knowledge-guided materials design and informatics-driven materials discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。