让大模型更懂材料科学,通过知识图谱引导渐进式训练
MELT: Materials-aware Continued Pre-training for Language Model Adaptation to Materials Science
- 基于科学文献构建材料知识图谱,设计由泛到专的训练课程
- 在多个材料科学任务上超越现有方法,显著提升实体表征能力
- 适合需要精准理解材料术语的研究者与工业应用
我们提出一种新型持续预训练方法MELT(MatEriaLs-aware continued pre-Training),专门用于高效适配预训练语言模型(PLMs)至材料科学领域。不同于以往仅关注构建领域语料的方法,MELT综合考虑语料与训练策略,因材料科学语料具有独特特征。为此,我们首先从科学文献中构建语义图谱,建立全面的材料知识库;利用该知识库,在适配过程中引入课程学习机制,从通用概念逐步过渡到专业术语。我们在多样化的基准上开展大量实验,验证MELT的有效性与普适性。综合评估有力支持其优势,表明其性能优于现有持续预训练方法。深入分析显示,相比现有适配方法,MELT能更有效表示材料实体,凸显其在广泛材料科学任务中的应用潜力。
原文摘要 · Abstract (English)
We introduce a novel continued pre-training method, MELT (MatEriaLs-aware continued pre-Training), specifically designed to efficiently adapt the pre-trained language models (PLMs) for materials science. Unlike previous adaptation strategies that solely focus on constructing domain-specific corpus, MELT comprehensively considers both the corpus and the training strategy, given that materials science corpus has distinct characteristics from other domains. To this end, we first construct a comprehensive materials knowledge base from the scientific corpus by building semantic graphs. Leveraging this extracted knowledge, we integrate a curriculum into the adaptation process that begins with familiar and generalized concepts and progressively moves toward more specialized terms. We conduct extensive experiments across diverse benchmarks to verify the effectiveness and generality of MELT. A comprehensive evaluation convincingly supports the strength of MELT, demonstrating superior performance compared to existing continued pre-training methods. The in-depth analysis also shows that MELT enables PLMs to effectively represent materials entities compared to the existing adaptation methods, thereby highlighting its broad applicability across a wide spectrum of materials science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。