arXiv:2509.19109cs.CL2025-09中稿 · TurkLang-2025 conf…被引 3

首个吉尔吉斯语命名实体识别数据集,助力低资源语言处理

Human-Annotated NER Dataset for the Kyrgyz Language

  • 构建了含39075个实体的吉尔吉斯语标注数据集
  • 多语言RoBERTa模型在该数据集上表现最优,精确率与召回率平衡良好
  • 适合从事低资源语言自然语言处理的研究者参考

我们提出首个针对吉尔吉斯语的人工标注命名实体识别数据集——KyrgyzNER。该数据集包含来自24.KG新闻门户的1,499篇新闻文章,共计10,900个句子和39,075个实体提及,覆盖27类命名实体。本文介绍了标注方案,讨论了标注过程中遇到的挑战,并提供了描述性统计。我们评估了多种命名实体识别模型,包括基于条件随机场的传统序列标注方法,以及在该数据集上微调的先进多语言Transformer模型。尽管所有模型在罕见实体类别上均表现不佳,但基于大规模多语言预训练的多语言RoBERTa变体展现出良好的精确率与召回率平衡。研究结果凸显了多语言预训练模型在资源有限语言处理中的潜力与挑战。尽管多语言RoBERTa表现最佳,其他多语言模型也取得了相近效果,表明未来探索更精细的标注方案可能为吉尔吉斯语处理流程评估提供更深入洞见。

原文摘要 · Abstract (English)

We introduce KyrgyzNER, the first manually annotated named entity recognition dataset for the Kyrgyz language. Comprising 1,499 news articles from the 24.KG news portal, the dataset contains 10,900 sentences and 39,075 entity mentions across 27 named entity classes. We show our annotation scheme, discuss the challenges encountered in the annotation process, and present the descriptive statistics. We also evaluate several named entity recognition models, including traditional sequence labeling approaches based on conditional random fields and state-of-the-art multilingual transformer-based models fine-tuned on our dataset. While all models show difficulties with rare entity categories, models such as the multilingual RoBERTa variant pretrained on a large corpus across many languages achieve a promising balance between precision and recall. These findings emphasize both the challenges and opportunities of using multilingual pretrained models for processing languages with limited resources. Although the multilingual RoBERTa model performed best, other multilingual models yielded comparable results. This suggests that future work exploring more granular annotation schemes may offer deeper insights for Kyrgyz language processing pipelines evaluation.

命名实体识别低资源语言多语言模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。