系统梳理北欧三国临床文本NLP研究现状,揭示语言间发展不均与资源共享不足问题。
Natural Language Processing for Electronic Health Records in Scandinavian Languages: Norwegian, Swedish, and Danish
- 系统回顾2010-2024年北欧三国临床文本NLP研究,覆盖挪威、瑞典、丹麦
- 瑞典研究占比64%(72篇),挪威18%(21篇),丹麦仅10%(11篇)
- 多语言研究少,模型资源与代码共享率低,尤其诺、丹两国进展滞后
背景:临床自然语言处理(NLP)利用计算方法提取、处理和分析非结构化临床文本数据,在多种临床任务中具有巨大潜力。目标:本研究开展系统综述,全面评估和分析大陆北欧语言(挪威语、瑞典语、丹麦语)的最新临床NLP方法。方法:2022年12月至2024年2月期间,在PubMed、ScienceDirect、Google Scholar、ACM数字图书馆和IEEE Xplore等数据库进行文献检索,并补充参考文献以完善搜索。最终纳入在2010–2024年间发表且使用英文描述、针对北欧主流语言临床文本的NLP研究论文。结果:共纳入113篇论文,其中18%(n=21)聚焦挪威语,64%(n=72)聚焦瑞典语,10%(n=11)聚焦丹麦语,8%(n=9)涉及多语言。总体上,该地区虽有积极进展,但语言间仍存在显著差距。基于Transformer的模型采用程度差异大;在去标识化等关键任务中,挪威语和丹麦语的研究远少于瑞典语。此外,数据、代码、预训练模型及迁移学习的共享率普遍偏低。结论:本综述全面评估了北欧电子健康记录(EHR)文本的临床NLP现状,指出了阻碍该领域快速发展的潜在障碍与挑战。
原文摘要 · Abstract (English)
Background: Clinical natural language processing (NLP) refers to the use of computational methods for extracting, processing, and analyzing unstructured clinical text data, and holds a huge potential to transform healthcare in various clinical tasks. Objective: The study aims to perform a systematic review to comprehensively assess and analyze the state-of-the-art NLP methods for the mainland Scandinavian clinical text. Method: A literature search was conducted in various online databases including PubMed, ScienceDirect, Google Scholar, ACM digital library, and IEEE Xplore between December 2022 and February 2024. Further, relevant references to the included articles were also used to solidify our search. The final pool includes articles that conducted clinical NLP in the mainland Scandinavian languages and were published in English between 2010 and 2024. Results: Out of the 113 articles, 18% (n=21) focus on Norwegian clinical text, 64% (n=72) on Swedish, 10% (n=11) on Danish, and 8% (n=9) focus on more than one language. Generally, the review identified positive developments across the region despite some observable gaps and disparities between the languages. There are substantial disparities in the level of adoption of transformer-based models. In essential tasks such as de-identification, there is significantly less research activity focusing on Norwegian and Danish compared to Swedish text. Further, the review identified a low level of sharing resources such as data, experimentation code, pre-trained models, and rate of adaptation and transfer learning in the region. Conclusion: The review presented a comprehensive assessment of the state-of-the-art Clinical NLP for electronic health records (EHR) text in mainland Scandinavian languages and, highlighted the potential barriers and challenges that hinder the rapid advancement of the field in the region.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。