专为挪威临床文本优化的BERT模型,提升医疗NLP表现
KliniskVestBERT: BERT Model Specialised to Norwegian Clinical Texts

- 在真实挪威临床文本上继续预训练现有BERT模型
- 在三个基准数据集和两个真实任务中均优于基线模型
- 适合需要处理挪威语医疗文本的研究与开发者
自然语言处理在医疗领域的应用日益广泛,亟需针对临床语言特性的语言模型。本文提出KliniskVestBERT,一套基于BERT的编码器模型,其在来自Helse Vest的大量真实去标识化挪威临床文本上进行预训练。我们对Nb-BERT-large、NorBERT3-large和ModernBERT三个现有模型进行了继续预训练。该数据集涵盖布克莫尔语和尼诺斯克语,包括出院小结、手术报告、护理记录等文档类型,覆盖挪威医疗场景的语言多样性。在三个合成临床基准数据集和两个真实世界任务上的评估显示,各临床专用模型均持续优于基线模型,证明领域特定预训练在临床NLP中的显著价值。该项目由Helse Vest各机构(Helse Bergen、Helse Fonna、Helse Førde、Helse Stavanger)与DIPS合作,由Helse Vest ICT牵头。
原文摘要 · Abstract (English)
The increasing application of Natural Language Processing (NLP) in healthcare demands language models specifically attuned to the complexities of clinical language. This work introduces KliniskVestBERT, a suite of three BERT-based encoder models pre-trained on a substantial corpus of real-world, de-identified Norwegian clinical texts from Helse Vest. We continue pretraining existing language models Nb-BERT-large, NorBERT3-large, and ModernBERT on our specialized clinical dataset. This dataset is based on a representative population of Helse Vest patients. The included document types are carefully curated to encompass a broad clinical spectrum in bokmål and nynorsk including discharge summaries, surgical reports, nursing notes etc. ensuring comprehensive representation of the linguistic landscape within Norwegian healthcare settings. Evaluation on three synthtetic Norwegian clinical benchmark datasets and two real-world problems demonstrates that each of our clinically specialized models consistently outperforms their baseline counterparts, highlighting the significant benefit of domain-specific pre-training for NLP tasks within the clinical domain. The project was a joint effort by all Helse Vest entities (Helse Bergen, Helse Fonna, Helse Førde and Helse Stavanger) with DIPS under the project lead of Helse Vest ICT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。