首个专为尼泊尔语设计的生成式大模型,填补语言空白
NepaliGPT: A Generative Language Model for the Nepali Language
- 基于自建德鲁纳加里语料库训练尼泊尔语生成模型
- 在4296组问答数据上达到26.3的困惑度和85.4%的因果一致性
- 为尼泊尔语NLP研究提供首个可用生成模型和评测基准
ChatGPT发布后,大型语言模型(LLMs)迅速流行,已涌现出数千种变体。然而,由于缺乏针对尼泊尔语的生成式语言模型,相关下游任务如微调尚未开展。为填补尼泊尔语自然语言处理领域的研究空白,本研究提出专为尼泊尔语定制的生成式大语言模型NepaliGPT。研究构建了一个来自多个来源的先进尼泊尔语文本语料库,命名为德鲁纳加里语料库(Devanagari Corpus)。同时,首次提出包含4,296个尼泊尔语问答对的NepaliGPT基准数据集。所提出的NepaliGPT在文本生成任务中表现优异:困惑度(Perplexity)为26.32245,ROUGE-1得分为0.2604,因果连贯性达81.25%,因果一致性为85.41%。
原文摘要 · Abstract (English)
After the release of ChatGPT, Large Language Models (LLMs) have gained huge popularity in recent days and thousands of variants of LLMs have been released. However, there is no generative language model for the Nepali language, due to which other downstream tasks, including fine-tuning, have not been explored yet. To fill this research gap in the Nepali NLP space, this research proposes \textit{NepaliGPT}, a generative large language model tailored specifically for the Nepali language. This research introduces an advanced corpus for the Nepali language collected from several sources, called the Devanagari Corpus. Likewise, the research introduces the first NepaliGPT benchmark dataset comprised of 4,296 question-answer pairs in the Nepali language. The proposed LLM NepaliGPT achieves the following metrics in text generation: Perplexity of 26.32245, ROUGE-1 score of 0.2604, causal coherence of 81.25\%, and causal consistency of 85.41\%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。