为尼泊尔语构建了首个大规模预训练模型,显著提升其文本理解与生成能力。
Development of Pre-Trained Transformer-based Models for the Nepali Language
- 基于27.5GB尼泊尔语数据,训练BERT、RoBERTa和GPT-2三类模型
- 在Nep-gLUE上达95.60分,比现有最佳模型高2分,生成任务也更优
- 首次探索单语尼泊尔语指令微调,为后续研究提供基础
基于Transformer的预训练语言模型在自然语言处理领域已占据主导地位。然而,全球约3200万人使用的尼泊尔语在此领域仍严重缺乏代表性,主要因缺乏单语语料库及可用资源有限。现有工作多集中于基础编码器模型,对解码器架构的研究则明显不足。为此,我们收集了27.5GB的尼泊尔语文本数据,约为此前最大语料库的2.4倍。基于此数据,我们为尼泊尔语独家预训练了BERT、RoBERTa和GPT-2三种模型,并进行了指令微调,探索其在单语尼泊尔语数据上的潜力,为未来研究奠定基础。我们的模型在Nep-gLUE基准测试中达到95.60分,优于现有最佳模型2分,同时在文本生成任务中也表现更优,证明了在理解与生成尼泊尔语方面的双重提升。
原文摘要 · Abstract (English)
Transformer-based pre-trained language models have dominated the field of Natural Language Processing (NLP) for quite some time now. However, the Nepali language, spoken by approximately 32 million people worldwide, remains significantly underrepresented in this domain. This underrepresentation is primarily attributed to the scarcity of monolingual data corpora and limited available resources for the Nepali language. While existing efforts have predominantly concentrated on basic encoder-based models, there is a notable gap in the exploration of decoder-based architectures. To address this gap, we have collected 27.5 GB of Nepali text data, approximately 2.4x larger than any previously available Nepali language corpus. Leveraging this data, we pre-trained three different models i.e., BERT, RoBERTa, and GPT-2, exclusively for the Nepali Language. Furthermore, we performed instruction tuning and explored its potential for monolingual Nepali data, providing a foundation for future research. Our models outperformed the existing best model by 2 points on Nep-gLUE benchmark, scoring 95.60 and also outperformed existing models on text generation tasks, demonstrating improvements in both understanding and generating Nepali text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。