构建首个大规模荷兰语医学语料库,支持医疗NLP研究。
Language corpora for the Dutch medical domain
- 整合翻译、筛选与开源资源构建语料
- 覆盖1.05亿文档,约360亿词元
- 免费开放于Hugging Face,适配预训练与下游任务
背景:荷兰语医学语料库稀缺,制约自然语言处理发展。方法:通过翻译英文数据集、从通用语料中识别医学文本、提取开放的荷兰语医学资源构建语料。结果:最终语料涵盖约1.05亿份文档,总计约360亿词元,已免费发布于Hugging Face。结论:本工作建立了首个可用于预训练及下游NLP任务的大规模荷兰语医学语言语料库。
原文摘要 · Abstract (English)
Background: Dutch medical corpora are scarce, limiting NLP development. Methods: We translated English datasets, identified medical text in generic corpora, and extracted open Dutch medical resources. Results: The resulting corpus comprises +- 36 billion tokens across the medical domain in about 105 million documents, freely available on Hugging Face. Conclusion: This work establishes the first large-scale Dutch medical language corpus for pre-training and downstream NLP tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。