用新标注框架从文献中提取材料制备-结构-性能关系,提升知识挖掘效率。
Structured Extraction of Process Structure Properties Relationships in Materials Science
- 设计专用标注方案,提取材料制备-结构-性能三元关系
- 微调LLM在高温材料领域表现优于基线模型,准确率提升显著
- 跨领域数据增强使传统模型媲美先进LLM,适合材料研究者使用
随着大语言模型(LLMs)的发展,海量学术论文中的非结构化文本日益可被用于材料发现,但挑战依然存在。尽管通用LLMs具备少样本和零样本学习能力,尤其在专家标注稀缺的材料领域极具价值,但未经适配的通用模型难以有效回答特定材料问题。为此,基于人工标注数据对LLMs进行微调成为实现结构化知识提取的关键。本研究提出一种新型标注框架,用于从科学文献中提取通用的工艺-结构-性能关系。我们利用128篇摘要构成的数据集,涵盖两个不同领域:高温材料(域I)与材料微观结构模拟中的不确定性量化(域II)。首先,基于领域特异性BERT变体MatBERT构建条件随机场(CRF)模型,并在域I上评估其性能;随后,在相同条件下对比该模型与微调后的GPT-4o模型。结果表明,微调后的LLM在域I上显著优于基线的BERT-CRF模型;然而,当引入域II的额外样本后,BERT-CRF模型性能达到与GPT-4o相当水平。这些发现验证了所提标注框架在结构化知识提取中的有效性,并凸显两种建模方法的互补优势。
原文摘要 · Abstract (English)
With the advent of large language models (LLMs), the vast unstructured text within millions of academic papers is increasingly accessible for materials discovery, although significant challenges remain. While LLMs offer promising few- and zero-shot learning capabilities, particularly valuable in the materials domain where expert annotations are scarce, general-purpose LLMs often fail to address key materials-specific queries without further adaptation. To bridge this gap, fine-tuning LLMs on human-labeled data is essential for effective structured knowledge extraction. In this study, we introduce a novel annotation schema designed to extract generic process-structure-properties relationships from scientific literature. We demonstrate the utility of this approach using a dataset of 128 abstracts, with annotations drawn from two distinct domains: high-temperature materials (Domain I) and uncertainty quantification in simulating materials microstructure (Domain II). Initially, we developed a conditional random field (CRF) model based on MatBERT, a domain-specific BERT variant, and evaluated its performance on Domain I. Subsequently, we compared this model with a fine-tuned LLM (GPT-4o from OpenAI) under identical conditions. Our results indicate that fine-tuning LLMs can significantly improve entity extraction performance over the BERT-CRF baseline on Domain I. However, when additional examples from Domain II were incorporated, the performance of the BERT-CRF model became comparable to that of the GPT-4o model. These findings underscore the potential of our schema for structured knowledge extraction and highlight the complementary strengths of both modeling approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。