arXiv:2501.11003cs.CL2025-01被引 6

为肯尼亚三种濒危语言构建开源语料库,助力非洲语言AI发展

Building low-resource African language corpora: A case study of Kidawida, Kalenjin and Dholuo

  • 通过众包收集母语者对话与朗读文本,构建平行与语音语料
  • 建成并开源三大语言语料库,支持模型训练与本地化应用开发
  • 推动非洲语言数字包容,适合本土开发者与语言保护研究者

自然语言处理是人工智能的关键领域,在公共卫生、农业、教育和商业中具有广泛应用。然而,由于缺乏充足的语言资源,许多非洲语言在数字化进程中仍被忽视。本文以肯尼亚三种低资源语言Kidaw'ida、Kalenjin和Dholuo为例,开展语料库建设的案例研究,旨在推进非洲社区的自然语言处理与语言学研究。项目历时一年,采用选择性众包方法,从母语者处收集文本与语音数据:(1) 记录对话并翻译成斯瓦希里语,形成双语语料;(2) 朗读书面文本以生成语音语料。所有资源已通过开放平台发布:平行文本语料上传至Zenodo,语音数据共享于Mozilla Common Voice,便于开发者持续贡献与使用。本项目展示了基层语料建设如何促进非洲语言融入人工智能创新。这些语料不仅填补资源空白,更推动语言多样性保护,赋能本地社区开发符合自身需求的NLP应用。随着肯尼亚等非洲国家加速数字化转型,建设本土语言资源对实现包容性增长至关重要。我们呼吁更多母语者与开发者参与共建与利用。

原文摘要 · Abstract (English)

Natural Language Processing is a crucial frontier in artificial intelligence, with broad applications in many areas, including public health, agriculture, education, and commerce. However, due to the lack of substantial linguistic resources, many African languages remain underrepresented in this digital transformation. This paper presents a case study on the development of linguistic corpora for three under-resourced Kenyan languages, Kidaw'ida, Kalenjin, and Dholuo, with the aim of advancing natural language processing and linguistic research in African communities. Our project, which lasted one year, employed a selective crowd-sourcing methodology to collect text and speech data from native speakers of these languages. Data collection involved (1) recording conversations and translation of the resulting text into Kiswahili, thereby creating parallel corpora, and (2) reading and recording written texts to generate speech corpora. We made these resources freely accessible via open-research platforms, namely Zenodo for the parallel text corpora and Mozilla Common Voice for the speech datasets, thus facilitating ongoing contributions and access for developers to train models and develop Natural Language Processing applications. The project demonstrates how grassroots efforts in corpus building can support the inclusion of African languages in artificial intelligence innovations. In addition to filling resource gaps, these corpora are vital in promoting linguistic diversity and empowering local communities by enabling Natural Language Processing applications tailored to their needs. As African countries like Kenya increasingly embrace digital transformation, developing indigenous language resources becomes essential for inclusive growth. We encourage continued collaboration from native speakers and developers to expand and utilize these corpora.

语料库非洲语言开源数据NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。