为缺乏资源的迈蒂利语构建了首个预训练语言模型maiBERT。
MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language
- 基于掩码语言建模,用自建语料库训练迈蒂利语BERT模型。
- 新闻分类任务中准确率达87.02%,优于NepBERTa和HindiBERT。
- 已开源,适合做情感分析、命名实体识别等下游任务。
低资源语言的自然语言理解仍是自然语言处理中的重大挑战,主要源于高质量数据和专用语言模型的匮乏。迈蒂利语虽有数百万使用者,却缺乏足够的计算资源,限制其在数字和人工智能应用中的使用。为此,我们提出了maiBERT,一个基于BERT、采用掩码语言建模(MLM)技术、专为迈蒂利语预训练的语言模型。该模型在新构建的迈蒂利语语料库上训练,并通过新闻分类任务进行评估。实验显示,maiBERT准确率达到87.02%,优于现有区域模型NepBERTa和HindiBERT,整体准确率提升0.13%,各类别性能提升5-7%。我们已在Hugging Face开源maiBERT,支持后续微调用于情感分析、命名实体识别(NER)等下游任务。
原文摘要 · Abstract (English)
Natural Language Understanding (NLU) for low-resource languages remains a major challenge in NLP due to the scarcity of high-quality data and language-specific models. Maithili, despite being spoken by millions, lacks adequate computational resources, limiting its inclusion in digital and AI-driven applications. To address this gap, we introducemaiBERT, a BERT-based language model pre-trained specifically for Maithili using the Masked Language Modeling (MLM) technique. Our model is trained on a newly constructed Maithili corpus and evaluated through a news classification task. In our experiments, maiBERT achieved an accuracy of 87.02%, outperforming existing regional models like NepBERTa and HindiBERT, with a 0.13% overall accuracy gain and 5-7% improvement across various classes. We have open-sourced maiBERT on Hugging Face enabling further fine-tuning for downstream tasks such as sentiment analysis and Named Entity Recognition (NER).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。