arXiv:2511.06531cs.CLcs.AI2025-11中稿 · IJCNLP-AACL被引 6

构建尼日利亚少数语言数据集,填补NLP研究空白

Ibom NLP: A Step Toward Inclusive Natural Language Processing for Nigeria's Minority Languages

  • 创建伊博州四语言文本数据集ibom,支持机器翻译与主题分类
  • 现有大模型在零样本和少样本下对这些语言表现差,但少样本可提升分类效果
  • 填补Google Translate与主流基准的空白,适合关注包容性NLP的研究者

尼日利亚人口超2亿,是非洲最人口稠密的国家,拥有超过500种语言,是全球语言多样性最高的国家之一。然而,自然语言处理(NLP)研究主要集中于豪萨语、伊博语、尼日利亚皮钦语和约鲁巴语这四种语言(占尼日利亚语言不足1%)。这主要由于这些语言缺乏可用于训练和应用NLP算法的文本数据。本文介绍ibom——一个面向尼日利亚沿海阿科伊博姆州四种语言(阿纳安语、埃菲克语、伊比比奥语、奥罗语)的机器翻译与主题分类数据集。这些语言未被纳入Google Translate或主要基准如Flores-200、SIB-200。我们拓展了Flores-200基准,并将翻译文本与基于SIB-200的主题标签对齐。评估表明,当前大语言模型在这些语言上的机器翻译表现不佳,无论在零样本还是少样本设置下。然而,随着少样本数量增加,主题分类性能稳步提升。

原文摘要 · Abstract (English)

Nigeria is the most populous country in Africa with a population of more than 200 million people. More than 500 languages are spoken in Nigeria and it is one of the most linguistically diverse countries in the world. Despite this, natural language processing (NLP) research has mostly focused on the following four languages: Hausa, Igbo, Nigerian-Pidgin, and Yoruba (i.e <1% of the languages spoken in Nigeria). This is in part due to the unavailability of textual data in these languages to train and apply NLP algorithms. In this work, we introduce ibom -- a dataset for machine translation and topic classification in four Coastal Nigerian languages from the Akwa Ibom State region: Anaang, Efik, Ibibio, and Oro. These languages are not represented in Google Translate or in major benchmarks such as Flores-200 or SIB-200. We focus on extending Flores-200 benchmark to these languages, and further align the translated texts with topic labels based on SIB-200 classification dataset. Our evaluation shows that current LLMs perform poorly on machine translation for these languages in both zero-and-few shot settings. However, we find the few-shot samples to steadily improve topic classification with more shots.

NLP多语言数据集包容性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。