BioClinical ModernBERT提升医学临床文本理解能力,支持长文本处理。
BioClinical ModernBERT: A State-of-the-Art Long-Context Encoder for Biomedical and Clinical NLP
- 基于535亿词元大规模语料持续预训练,适配医学临床领域
- 在4项下游任务中超越现有模型,尤其擅长长文本分析
- 提供大小两个版本及训练检查点,适合研究与应用开发
基于Transformer的编码器模型在生物医学和临床自然语言处理中至关重要,其双向自注意力机制能高效提取非结构化文本中的结构化信息。然而,相比解码器模型,编码器发展较慢,导致在生物医学和临床场景中领域适应性有限。本文提出BioClinical ModernBERT,一种基于最新ModernBERT改进的领域适配编码器,具备长文本处理能力,并在速度与性能上实现显著提升。该模型通过在迄今最大的生物医学与临床语料库(超53.5亿词元)上进行持续预训练构建,克服了以往临床编码器依赖单一来源数据的局限性,整合来自20个不同机构、领域与地理区域的数据。在涵盖多种应用场景的四项下游任务中,该模型表现优于现有生物医学与临床编码器。我们发布了基础版(150M参数)和大型版(396M参数)模型,以及训练检查点以支持进一步研究。
原文摘要 · Abstract (English)
Encoder-based transformer models are central to biomedical and clinical Natural Language Processing (NLP), as their bidirectional self-attention makes them well-suited for efficiently extracting structured information from unstructured text through discriminative tasks. However, encoders have seen slower development compared to decoder models, leading to limited domain adaptation in biomedical and clinical settings. We introduce BioClinical ModernBERT, a domain-adapted encoder that builds on the recent ModernBERT release, incorporating long-context processing and substantial improvements in speed and performance for biomedical and clinical NLP. BioClinical ModernBERT is developed through continued pretraining on the largest biomedical and clinical corpus to date, with over 53.5 billion tokens, and addresses a key limitation of prior clinical encoders by leveraging 20 datasets from diverse institutions, domains, and geographic regions, rather than relying on data from a single source. It outperforms existing biomedical and clinical encoders on four downstream tasks spanning a broad range of use cases. We release both base (150M parameters) and large (396M parameters) versions of BioClinical ModernBERT, along with training checkpoints to support further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。