arXiv:2503.05373cs.CL2025-03被引 4

利用医学语义类型依赖关系提升临床实体识别准确率

Leveraging Semantic Type Dependencies for Clinical Named Entity Recognition

  • 引入医学术语间的语义依赖关系作为额外特征
  • 在多个数据集上显著提升命名实体识别效果
  • 首次实现单次处理多于三种依赖关系的编码方法

以往临床关系抽取研究利用临床知识库中的语义类型信息作为实体表示的一部分。本文进一步挖掘领域特异性语义类型依赖关系,编码句子中匹配统一医学语言系统(UMLS)概念的词元段与其他词元之间的关系。我们实现了该方法,并在不同命名实体识别架构(如BiLSTM-CRF和BiLSTM-GCN-CRF)上,使用多种预训练临床嵌入模型(如BERT、BioBERT、UMLSBert)进行对比实验。在临床数据集上的实验结果表明,利用领域特定的语义类型依赖可显著提升某些情况下的命名实体识别效果。本工作也是首个在单次处理中利用超过三种依赖关系编码的NER研究。

原文摘要 · Abstract (English)

Previous work on clinical relation extraction from free-text sentences leveraged information about semantic types from clinical knowledge bases as a part of entity representations. In this paper, we exploit additional evidence by also making use of domain-specific semantic type dependencies. We encode the relation between a span of tokens matching a Unified Medical Language System (UMLS) concept and other tokens in the sentence. We implement our method and compare against different named entity recognition (NER) architectures (i.e., BiLSTM-CRF and BiLSTM-GCN-CRF) using different pre-trained clinical embeddings (i.e., BERT, BioBERT, UMLSBert). Our experimental results on clinical datasets show that in some cases NER effectiveness can be significantly improved by making use of domain-specific semantic type dependencies. Our work is also the first study generating a matrix encoding to make use of more than three dependencies in one pass for the NER task.

临床NLP实体识别语义依赖UMLS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。