发布首个大规模西班牙语临床NLP数据集与模型,助力医疗文本分析
ClinText-SP and RigoBERTa Clinical: a new set of open resources for Spanish Clinical NLP
- 构建了最大公开西班牙语临床语料库ClinText-SP,涵盖多源医学文献
- RigoBERTa Clinical在多个临床任务上超越现有模型性能
- 适合从事西班牙语医疗文本分析的研究者和开发者
我们提出西班牙语临床自然语言处理的新资源:最大规模的公开临床语料库ClinText-SP,以及基于该语料库训练的先进临床编码模型RigoBERTa Clinical。语料库通过整合来自医学期刊的临床案例和共享任务标注数据精心构建,内容丰富且来源多样。RigoBERTa Clinical通过在该综合数据集上进行领域自适应预训练,显著优于现有模型,在多个临床NLP基准测试中表现突出。我们公开发布该数据集与模型,旨在为研究社区提供强大工具,推动临床NLP发展,助力医疗应用进步。
原文摘要 · Abstract (English)
We present a novel contribution to Spanish clinical natural language processing by introducing the largest publicly available clinical corpus, ClinText-SP, along with a state-of-the-art clinical encoder language model, RigoBERTa Clinical. Our corpus was meticulously curated from diverse open sources, including clinical cases from medical journals and annotated corpora from shared tasks, providing a rich and diverse dataset that was previously difficult to access. RigoBERTa Clinical, developed through domain-adaptive pretraining on this comprehensive dataset, significantly outperforms existing models on multiple clinical NLP benchmarks. By publicly releasing both the dataset and the model, we aim to empower the research community with robust resources that can drive further advancements in clinical NLP and ultimately contribute to improved healthcare applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。