多语言医学BERT通过领域适配提升低资源语言任务表现。
Multilingual BERT language model for medical tasks: Evaluation on domain-specific adaptation and cross-linguality
- 在荷兰语、罗马尼亚语和西班牙语临床文本上进行医学领域微调。
- 临床领域适配模型性能优于通用生物医学模型,提升显著。
- 验证了跨语言迁移能力,适合低资源医疗NLP系统开发。
在多语言医疗应用中,低资源语言的领域专用自然语言处理工具仍十分有限。尽管多语言BERT有望缓解语言差距,但其在低资源语言上的医学任务研究仍不充分。本研究探讨在特定领域语料上进一步预训练对医学任务性能的影响,聚焦荷兰语、罗马尼亚语和西班牙语。通过四项实验构建医学领域模型,并在三项下游任务上进行微调:荷兰语临床笔记的自动患者筛查、罗马尼亚语与西班牙语临床笔记的命名实体识别。结果表明,领域适配显著提升了任务性能;进一步区分临床与通用生物医学领域后,临床领域适配模型表现优于通用生物医学模型。同时观察到跨语言迁移的证据。研究揭示了领域适配与跨语言能力在医学NLP中的可行性,为低资源语言场景下构建多语言医疗NLP系统提供了重要指导。
原文摘要 · Abstract (English)
In multilingual healthcare applications, the availability of domain-specific natural language processing(NLP) tools is limited, especially for low-resource languages. Although multilingual bidirectional encoder representations from transformers (BERT) offers a promising motivation to mitigate the language gap, the medical NLP tasks in low-resource languages are still underexplored. Therefore, this study investigates how further pre-training on domain-specific corpora affects model performance on medical tasks, focusing on three languages: Dutch, Romanian and Spanish. In terms of further pre-training, we conducted four experiments to create medical domain models. Then, these models were fine-tuned on three downstream tasks: Automated patient screening in Dutch clinical notes, named entity recognition in Romanian and Spanish clinical notes. Results show that domain adaptation significantly enhanced task performance. Furthermore, further differentiation of domains, e.g. clinical and general biomedical domains, resulted in diverse performances. The clinical domain-adapted model outperformed the more general biomedical domain-adapted model. Moreover, we observed evidence of cross-lingual transferability. Moreover, we also conducted further investigations to explore potential reasons contributing to these performance differences. These findings highlight the feasibility of domain adaptation and cross-lingual ability in medical NLP. Within the low-resource language settings, these findings can provide meaningful guidance for developing multilingual medical NLP systems to mitigate the lack of training data and thereby improve the model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。