用轻量模型识别等离子体论文中的嵌套实体,提升科研文献检索效率。
Nested Named Entity Recognition in Plasma Physics Research Articles
- 基于BERT-CRF构建分类型实体识别模型,支持嵌套结构。
- 自建16类标注语料库,覆盖等离子体物理专业术语。
- 通过超参优化与模型特化,显著提升复杂文本识别准确率。
命名实体识别(NER)是自然语言处理中从非结构化文本中识别并提取关键实体的重要任务。本文首次将NER应用于等离子体物理研究文章,针对该领域文本高度复杂、上下文丰富等特点,提出一种轻量级的嵌套实体识别方法。首先,我们构建了一个包含16个类别、专为嵌套NER设计的等离子体物理语料库。其次,采用特定实体类型的模型专业化策略,训练独立的BERT-CRF模型以识别不同类型的实体。第三,引入系统性超参数优化流程,进一步提升模型性能。本工作推动了等离子体物理领域实体识别技术的发展,为研究人员高效导航和分析科学文献提供了坚实基础。
原文摘要 · Abstract (English)
Named Entity Recognition (NER) is an important task in natural language processing that aims to identify and extract key entities from unstructured text. We present a novel application of NER in plasma physics research articles and address the challenges of extracting specialized entities from scientific text in this domain. Research articles in plasma physics often contain highly complex and context-rich content that must be extracted to enable, e.g., advanced search. We propose a lightweight approach based on encoder-transformers and conditional random fields to extract (nested) named entities from plasma physics research articles. First, we annotate a plasma physics corpus with 16 classes specifically designed for the nested NER task. Second, we evaluate an entity-specific model specialization approach, where independent BERT-CRF models are trained to recognize individual entity types in plasma physics text. Third, we integrate an optimization process to systematically fine-tune hyperparameters and enhance model performance. Our work contributes to the advancement of entity recognition in plasma physics and also provides a foundation to support researchers in navigating and analyzing scientific literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。