arXiv:2606.03250cs.CL2026-06被引 3

针对德语医学文本,训练出更精准的专用语言模型。

The Word and the Way: Strategies for Domain-Specific BERT Pre-Training in German Medical NLP

论文配图:The Word and the Way: Strategies for Domain-Specific BERT Pre-Training in German Medical NLP
图 1 · 摘自论文原文
  • 用13.5GB医学数据从头训练德语RoBERTa模型。
  • 在五项任务中四项超越现有模型,达到新纪录。
  • 特定场景下从头训练比微调更有效,适合专业文本。

数字医疗产生大量临床文本,可用于AI辅助应用,但德语生物医学语言模型仍受限于旧架构或有限数据。我们提出ChristBERT(Clinical- and Healthcare-Related Issues and Subjects Tuned BERT),基于RoBERTa的德语领域专用语言模型家族,使用13.5GB科学论文、临床文本、健康类网络内容及翻译临床资源进行训练。为探究德语临床NLP中领域适配策略的影响,我们对比了继续预训练、从头训练和领域特定词汇适应。模型在三项医学命名实体识别任务和两项文本分类任务上评估。ChristBERT在五项基准中的四项表现优于现有通用及医学德语模型,建立德语临床语言建模新基准。结果表明最优策略取决于任务:在高度专业化的临床文本中,从头训练效果更佳;而在较常见的医学文本上,继续预训练表现更优。所有模型已公开发布,以支持未来德语医学NLP研究与应用。

原文摘要 · Abstract (English)

Digital healthcare generates vast amounts of clinical text that can support AI-assisted applications, yet German biomedical language models remain limited by older architectures or restricted training data. We present ChristBERT (Clinical- and Healthcare-Related Issues and Subjects Tuned BERT), a family of domain-specific German RoBERTa-based language models trained on a 13.5GB corpus of scientific publications, clinical texts, health-related web content, and translated clinical resources. To investigate the impact of domain adaptation strategies in German clinical NLP, we compare continued pre-training, training from scratch, and domain-specific vocabulary adaptation. The resulting models are evaluated on three medical named entity recognition tasks and two text classification tasks. ChristBERT consistently outperforms existing general-purpose and medical German language models on four of five benchmarks and establishes a new state of the art for German clinical language modeling. Our results show that the optimal adaptation strategy is task-dependent: in our evaluation, training from scratch is particularly effective for highly specialized clinical texts, whereas continued pre-training performs well on more commonly written medical texts. All models are publicly released to support future research and applications in German medical NLP.

医学NLPBERT德语领域适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。