提出多BERT集成模型,提升孟加拉语医学实体识别准确率至89.58%。
Bangla MedER: Multi-BERT Ensemble Approach for the Recognition of Bangla Medical Entity
- 采用多BERT集成策略,融合Bert、DistilBERT等模型优势
- 在自建数据集上达到89.58%准确率,较单BERT提升11.80%
- 为低资源语言医学NLP提供可复用的模型与数据基础
医学实体识别(MedER)是提取医学文本中关键信息的重要自然语言处理任务。尽管英文领域的研究已较为成熟,但像孟加拉语这类低资源语言仍缺乏系统研究。本文首先评估了BERT、DistilBERT、ELECTRA和RoBERTa等多种Transformer模型在孟加拉语医学文本中的表现,并提出一种新型多BERT集成方法。该方法在自建的高质量孟加拉语医学实体数据集上取得了89.58%的最高准确率,相比单层BERT模型提升11.80%,显著优于所有基线模型。研究还针对低资源语言标注数据稀缺的问题,构建了专门用于孟加拉语医学生物实体识别的数据集。实验结果验证了所提模型的鲁棒性与实用性,为低资源语言医学NLP的发展提供了坚实基础。
原文摘要 · Abstract (English)
Medical Entity Recognition (MedER) is an essential NLP task for extracting meaningful entities from the medical corpus. Nowadays, MedER-based research outcomes can remarkably contribute to the development of automated systems in the medical sector, ultimately enhancing patient care and outcomes. While extensive research has been conducted on MedER in English, low-resource languages like Bangla remain underexplored. Our work aims to bridge this gap. For Bangla medical entity recognition, this study first examined a number of transformer models, including BERT, DistilBERT, ELECTRA, and RoBERTa. We also propose a novel Multi-BERT Ensemble approach that outperformed all baseline models with the highest accuracy of 89.58%. Notably, it provides an 11.80% accuracy improvement over the single-layer BERT model, demonstrating its effectiveness for this task. A major challenge in MedER for low-resource languages is the lack of annotated datasets. To address this issue, we developed a high-quality dataset tailored for the Bangla MedER task. The dataset was used to evaluate the effectiveness of our model through multiple performance metrics, demonstrating its robustness and applicability. Our findings highlight the potential of Multi-BERT Ensemble models in improving MedER for Bangla and set the foundation for further advancements in low-resource medical NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。