arXiv:2511.20182cs.CL2025-11

首个面向吉尔吉斯语的高效BERT模型,助力低资源语言NLP发展

KyrgyzBERT: A Compact, Efficient Language Model for Kyrgyz NLP

  • 基于吉尔吉斯语形态结构设计专用分词器,模型仅35.9M参数
  • 在自建情感分析数据集上达0.8280的F1分数,性能超越五倍大的mBERT
  • 开源全部模型、数据与代码,推动吉尔吉斯语NLP研究

吉尔吉斯语作为低资源语言,缺乏基础NLP工具。为此,我们提出KyrgyzBERT,首个公开的吉尔吉斯语单语BERT模型,包含35.9M参数,并采用针对该语言形态结构定制的分词器。为评估性能,我们构建了kyrgyz-sst2,一个通过翻译Stanford Sentiment Treebank并人工标注完整测试集形成的情感分析基准。在该数据集上,经微调的KyrgyzBERT达到0.8280的F1分数,表现媲美参数量五倍于它的mBERT模型。所有模型、数据及代码均已开源,以支持未来吉尔吉斯语NLP研究。

原文摘要 · Abstract (English)

Kyrgyz remains a low-resource language with limited foundational NLP tools. To address this gap, we introduce KyrgyzBERT, the first publicly available monolingual BERT-based language model for Kyrgyz. The model has 35.9M parameters and uses a custom tokenizer designed for the language's morphological structure. To evaluate performance, we create kyrgyz-sst2, a sentiment analysis benchmark built by translating the Stanford Sentiment Treebank and manually annotating the full test set. KyrgyzBERT fine-tuned on this dataset achieves an F1-score of 0.8280, competitive with a fine-tuned mBERT model five times larger. All models, data, and code are released to support future research in Kyrgyz NLP.

语言模型低资源语言BERT吉尔吉斯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。