arXiv:2507.18542cs.CL2025-07ACL被引 3

解决生物医学实体识别中的嵌套与标注不一致问题

Effective Multi-Task Learning for Biomedical Named Entity Recognition

  • 采用基于槽的循环单元结构,动态调整损失函数以适应不同数据集
  • 在跨语料库评估中表现优异,提升跨领域泛化能力
  • 适合需要处理多源生物医学文本的科研人员与医疗AI开发者

生物医学命名实体识别因术语复杂和数据集标注不一致而面临挑战。本文提出SRU-NER(基于槽的循环单元命名实体识别),一种可处理嵌套实体并利用有效多任务学习整合多数据集的新方法。通过动态调整损失计算,避免对目标数据集中不存在的实体类型进行惩罚,缓解标注缺失问题。大量实验包括跨语料库评估和人工判断显示,SRU-NER在生物医学及通用领域命名实体识别任务中表现具有竞争力,并显著提升跨领域泛化能力。

原文摘要 · Abstract (English)

Biomedical Named Entity Recognition presents significant challenges due to the complexity of biomedical terminology and inconsistencies in annotation across datasets. This paper introduces SRU-NER (Slot-based Recurrent Unit NER), a novel approach designed to handle nested named entities while integrating multiple datasets through an effective multi-task learning strategy. SRU-NER mitigates annotation gaps by dynamically adjusting loss computation to avoid penalizing predictions of entity types absent in a given dataset. Through extensive experiments, including a cross-corpus evaluation and human assessment of the model's predictions, SRU-NER achieves competitive performance in biomedical and general-domain NER tasks, while improving cross-domain generalization.

命名实体识别多任务学习生物医学文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。