用AI从文献中自动提取燃料电池催化剂信息,准确率超80%。
Extracting ORR Catalyst Information for Fuel Cell from Scientific Literature
- 用BERT模型做命名实体和关系抽取,专用于催化剂文献分析。
- 最佳模型达82.19%的识别准确率,关系抽取也超66%。
- 适合材料科学、能源领域研究者快速获取文献关键数据。
氧还原反应(ORR)催化剂对提升燃料电池效率至关重要,是材料科学研究的核心。然而,从海量科学文献中提取结构化催化剂信息仍面临挑战。本文提出基于DyGIE++框架的命名实体识别(NER)与关系抽取(RE)方法,结合MatSciBERT和PubMedBERT等多种预训练BERT模型,从文献中提取ORR催化剂相关数据,并构建了面向材料信息学的燃料电池语料库(FC-CoMIcs)。通过人工标注12类关键实体及两类关系,完成数据标注、整合与模型微调,以提升信息提取精度。实验表明,微调后的PubMedBERT在NER上达到82.19%的F1分数,MatSciBERT在RE上取得66.10%的最佳表现。与人工标注对比显示,微调模型具备可靠性和可扩展性,适用于大规模自动化文献分析。结果还表明,领域专用的BERT模型(如PubMedBERT)在该任务中优于通用科学模型(如BlueBERT)。
原文摘要 · Abstract (English)
The oxygen reduction reaction (ORR) catalyst plays a critical role in enhancing fuel cell efficiency, making it a key focus in material science research. However, extracting structured information about ORR catalysts from vast scientific literature remains a significant challenge due to the complexity and diversity of textual data. In this study, we propose a named entity recognition (NER) and relation extraction (RE) approach using DyGIE++ with multiple pre-trained BERT variants, including MatSciBERT and PubMedBERT, to extract ORR catalyst-related information from the scientific literature, which is compiled into a fuel cell corpus for materials informatics (FC-CoMIcs). A comprehensive dataset was constructed manually by identifying 12 critical entities and two relationship types between pairs of the entities. Our methodology involves data annotation, integration, and fine-tuning of transformer-based models to enhance information extraction accuracy. We assess the impact of different BERT variants on extraction performance and investigate the effects of annotation consistency. Experimental evaluations demonstrate that the fine-tuned PubMedBERT model achieves the highest NER F1-score of 82.19% and the MatSciBERT model attains the best RE F1-score of 66.10%. Furthermore, the comparison with human annotators highlights the reliability of fine-tuned models for ORR catalyst extraction, demonstrating their potential for scalable and automated literature analysis. The results indicate that domain-specific BERT models outperform general scientific models like BlueBERT for ORR catalyst extraction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。