解决德语患者文本医学概念归一化的低资源难题
Medical Concept Normalization in a Low-Resource Setting
- 构建德语医疗论坛语料库并标注UMLS概念
- 多语言Transformer模型优于传统字符串匹配方法
- 揭示常见错误模式并提出改进方向
在生物医学自然语言处理领域,医学概念归一化是将概念提及准确映射到大型知识库的关键任务。然而,在数据与资源有限的低资源环境下,该任务更具挑战性。本文研究德语非专业文本中医学概念归一化的难点。由于缺乏合适数据集,作者基于德语医疗在线论坛帖子构建了标注数据集,使用统一医学语言系统(UMLS)进行概念标注。实验表明,基于多语言Transformer的模型显著优于字符串相似性方法。同时考察了上下文信息对普通用户提及归一化的作用,但结果反而更差。基于表现最佳模型,本文进行了系统性误差分析,并提出缓解常见错误的潜在改进策略。
原文摘要 · Abstract (English)
In the field of biomedical natural language processing, medical concept normalization is a crucial task for accurately mapping mentions of concepts to a large knowledge base. However, this task becomes even more challenging in low-resource settings, where limited data and resources are available. In this thesis, I explore the challenges of medical concept normalization in a low-resource setting. Specifically, I investigate the shortcomings of current medical concept normalization methods applied to German lay texts. Since there is no suitable dataset available, a dataset consisting of posts from a German medical online forum is annotated with concepts from the Unified Medical Language System. The experiments demonstrate that multilingual Transformer-based models are able to outperform string similarity methods. The use of contextual information to improve the normalization of lay mentions is also examined, but led to inferior results. Based on the results of the best performing model, I present a systematic error analysis and lay out potential improvements to mitigate frequent errors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。