用语言模型提取材料名隐含信息,提升性质预测准确率
From Tokens to Materials: Leveraging Language Models for Scientific Discovery
- 用MatBERT等领域模型提取材料名中的隐含知识
- 第三层嵌入结合上下文平均法效果最佳
- 专用分词技术对保持分子式完整很关键
探索语言模型在材料科学中的预测能力是当前研究热点。本研究考察了语言模型嵌入在材料性质预测中的应用,对比了多种上下文嵌入方法及预训练模型,包括BERT和GPT。结果表明,领域专用模型如MatBERT显著优于通用模型,能更有效地从化合物名称和材料性质中提取隐含知识。研究发现,MatBERT第三层的信息密集型嵌入,配合上下文平均策略,是最优的材料-性质关系捕捉方式。同时识别出关键的“分词器效应”,强调需采用专用文本处理技术以保留完整化合物名称并维持一致的词元数量。这些发现凸显了领域专用训练与分词在材料科学应用中的价值,为通过人工智能加速新材料发现提供了可行路径。
原文摘要 · Abstract (English)
Exploring the predictive capabilities of language models in material science is an ongoing interest. This study investigates the application of language model embeddings to enhance material property prediction in materials science. By evaluating various contextual embedding methods and pre-trained models, including Bidirectional Encoder Representations from Transformers (BERT) and Generative Pre-trained Transformers (GPT), we demonstrate that domain-specific models, particularly MatBERT significantly outperform general-purpose models in extracting implicit knowledge from compound names and material properties. Our findings reveal that information-dense embeddings from the third layer of MatBERT, combined with a context-averaging approach, offer the most effective method for capturing material-property relationships from the scientific literature. We also identify a crucial "tokenizer effect," highlighting the importance of specialized text processing techniques that preserve complete compound names while maintaining consistent token counts. These insights underscore the value of domain-specific training and tokenization in materials science applications and offer a promising pathway for accelerating the discovery and development of new materials through AI-driven approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。