提升罕见病表型识别与标准化准确率,解决医学文本表述多样难题。
KU AIGEN ICL EDI@BC8 Track 3: Advancing Phenotype Named Entity Recognition and Normalization for Dysmorphology Physical Examination Reports
- 采用命名实体识别结合同义词替换增强数据,提升表型提取效果。
- 在精确提取与标准化上比平均分高2.6%,标准化得分高出1.9%。
- 适合医疗信息抽取、生物医学知识图谱构建等研究者参考。
BioCreative8 Track 3旨在从电子健康记录(EHR)文本中提取表型医学发现,并将其归一化为人类表型本体(HPO)术语。由于表型描述存在多种表达形式,准确归一化面临挑战。为此,我们测试了多种命名实体识别模型,并引入同义词边缘化等数据增强技术优化归一化步骤。所提流程在精确提取与归一化上的F1分数比所有参赛方案的均值高出2.6%,归一化F1分数超出均值1.9%。该成果推动了自动化医学数据提取与归一化技术发展,为生物医学领域未来研究与应用提供可行路径。
原文摘要 · Abstract (English)
The objective of BioCreative8 Track 3 is to extract phenotypic key medical findings embedded within EHR texts and subsequently normalize these findings to their Human Phenotype Ontology (HPO) terms. However, the presence of diverse surface forms in phenotypic findings makes it challenging to accurately normalize them to the correct HPO terms. To address this challenge, we explored various models for named entity recognition and implemented data augmentation techniques such as synonym marginalization to enhance the normalization step. Our pipeline resulted in an exact extraction and normalization F1 score 2.6\% higher than the mean score of all submissions received in response to the challenge. Furthermore, in terms of the normalization F1 score, our approach surpassed the average performance by 1.9\%. These findings contribute to the advancement of automated medical data extraction and normalization techniques, showcasing potential pathways for future research and application in the biomedical domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。