arXiv:2505.05864cs.CL2025-05

用符号标记提升材料科学文本挖掘准确率

Symbol-based entity marker highlighting for enhanced text mining in materials science with generative AI

  • 结合多步与直接法优势,先识别实体再结构化数据
  • 在三个数据集上实体识别F1提升最高达58%,关系识别提升83%
  • 适合需要高精度提取材料数据的研究者使用

构建实验数据集对于推动数据驱动的科学发现至关重要。近年来自然语言处理(NLP)的进步使得从非结构化科学文献中自动提取结构化数据成为可能。现有方法——多步法与直接法——虽具价值,但独立应用时各有局限。本文提出一种新型混合文本挖掘框架,融合两类方法优势,将非结构化科学文本转化为结构化数据:首先将原始文本转为带实体标注的文本,再进一步结构化。此外,我们引入一种简单有效的实体标记技术,通过符号注释突出目标实体,显著提升实体识别性能。该基于实体标记的混合方法在三个基准数据集(MatScholar、SOFC、SOFC slot NER)上持续优于先前方法,实体级F1分数最高提升58%,关系级F1分数最高提升83%。

原文摘要 · Abstract (English)

The construction of experimental datasets is essential for expanding the scope of data-driven scientific discovery. Recent advances in natural language processing (NLP) have facilitated automatic extraction of structured data from unstructured scientific literature. While existing approaches-multi-step and direct methods-offer valuable capabilities, they also come with limitations when applied independently. Here, we propose a novel hybrid text-mining framework that integrates the advantages of both methods to convert unstructured scientific text into structured data. Our approach first transforms raw text into entity-recognized text, and subsequently into structured form. Furthermore, beyond the overall data structuring framework, we also enhance entity recognition performance by introducing an entity marker-a simple yet effective technique that uses symbolic annotations to highlight target entities. Specifically, our entity marker-based hybrid approach not only consistently outperforms previous entity recognition approaches across three benchmark datasets (MatScholar, SOFC, and SOFC slot NER) but also improve the quality of final structured data-yielding up to a 58% improvement in entity-level F1 score and up to 83% improvement in relation-level F1 score compared to direct approach.

文本挖掘材料科学NLP实体识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。