构建首个医学乌尔都语命名实体识别基准数据集
BioUNER: A Benchmark Dataset for Clinical Urdu Named Entity Recognition
- 从新闻、处方和医院博客爬取数据,三名医学背景标注员标注15.3万词元
- 标注一致性达0.78,验证数据集为金标准;XLM-RoBERTa表现最佳
- 为乌尔都语医学自然语言处理提供关键资源,适合多语言医疗研究者
本文提出首个生物医学乌尔都语命名实体识别(BioUNER)金标准基准数据集,通过爬取在线乌尔都语新闻门户、医疗处方及医院健康博客与网站获取数据。经预处理后,三位具备医学背景的母语标注员使用Doccano工具对15.3万词元进行标注。标注后,内在与外在评估显示标注者间一致性达0.78,证实数据集质量可靠。为验证数据集实用性,我们测试了支持向量机(SVM)、长短期记忆网络(LSTM)、多语言BERT(mBERT)及XLM-RoBERTa等模型。结果表明,该数据集可作为乌尔都语医学自然语言处理的可靠基准,丰富了乌尔都语语言资源。
原文摘要 · Abstract (English)
In this article, we present a gold-standard benchmark dataset for Biomedical Urdu Named Entity Recognition (BioUNER), developed by crawling health-related articles from online Urdu news portals, medical prescriptions, and hospital health blogs and websites. After preprocessing, three native annotators with familiarity in the medical domain participated in the annotation process using the Doccano text annotation tool and annotated 153K tokens. Following annotation, the proposed BioiUNER dataset was evaluated both intrinsically and extrinsically. An inter-annotator agreement score of 0.78 was achieved, thereby validating the dataset as gold-standard quality. To demonstrate the utility and benchmarking capability of the dataset, we evaluated several machine learning and deep learning models, including Support Vector Machines (SVM), Long Short-Term Memory networks (LSTM), Multilingual BERT (mBERT), and XLM-RoBERTa. The gold-standard BioUNER dataset serves as a reliable benchmark and a valuable addition to Urdu language processing resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。