提升合金文献中多组元属性提取精度,融合指针网络与注意力机制。
Enhanced Multi-Tuple Extraction for Alloys: Integrating Pointer Networks and Augmented Attention
- 用MatSciBERT结合指针网络和实体间/内注意力建模复杂关系。
- 在含1至4个属性的数据集上F1达0.963~0.753,随机数据集0.854。
- 适合材料领域科研人员高效获取结构化实验数据。
从科学文献中提取高质量结构化信息对推动数据驱动材料设计至关重要。尽管自然语言处理在数据集提取方面已有大量研究,但因元组间复杂关联与上下文模糊性,科学文献中的多组元提取方法仍稀缺。本文针对多主元素合金的力学性能提取,提出一种新框架:结合基于MatSciBERT的实体抽取模型、指针网络以及利用实体间与实体内注意力的分配模型。在多组元提取上的严格实验表明,该模型在含1、2、3、4个元组的数据集上分别取得0.963、0.947、0.848和0.753的F1分数,随机筛选数据集亦达0.854。结果验证了模型在精准生成结构化信息方面的强大能力,为大型语言模型提供可靠替代方案,助力研究人员实现数据驱动创新。
原文摘要 · Abstract (English)
Extracting high-quality structured information from scientific literature is crucial for advancing material design through data-driven methods. Despite the considerable research in natural language processing for dataset extraction, effective approaches for multi-tuple extraction in scientific literature remain scarce due to the complex interrelations of tuples and contextual ambiguities. In the study, we illustrate the multi-tuple extraction of mechanical properties from multi-principal-element alloys and presents a novel framework that combines an entity extraction model based on MatSciBERT with pointer networks and an allocation model utilizing inter- and intra-entity attention. Our rigorous experiments on tuple extraction demonstrate impressive F1 scores of 0.963, 0.947, 0.848, and 0.753 across datasets with 1, 2, 3, and 4 tuples, confirming the effectiveness of the model. Furthermore, an F1 score of 0.854 was achieved on a randomly curated dataset. These results highlight the model's capacity to deliver precise and structured information, offering a robust alternative to large language models and equipping researchers with essential data for fostering data-driven innovations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。