arXiv:2502.17364cs.CLcs.AI2025-02综述被引 9

系统梳理十年来约鲁巴语NLP研究,揭示资源与技术瓶颈。

Bridging Gaps in Natural Language Processing for Yorùbá: A Systematic Review of a Decade of Progress and Prospects

  • 基于105篇论文的系统综述,分析约鲁巴语NLP发展脉络
  • 发现标注语料稀缺、预训练模型匮乏、声调复杂等核心挑战
  • 适合关注非洲语言AI、低资源语言研究者参考

自然语言处理(NLP)作为人工智能的重要分支,正日益成为机器理解人类语言的关键。尽管社交媒体等平台每日产生大量数据,推动了NLP广泛应用,但多数非洲语言仍面临资源匮乏等困境。约鲁巴语作为一种声调丰富、形态复杂的非洲语言,其NLP发展也受限。本文通过系统文献综述,全面分析2014至2024年间105篇高质量研究,聚焦约鲁巴语NLP的资源、技术与应用。研究指出,标注语料稀少、预训练模型缺失及声调与变音符号依赖是主要障碍。尽管规则方法仍占主导,近年已出现多语言与单语资源增长趋势。然而,语言混用与数字使用中语言流失等社会文化因素仍制约发展。本综述整合现有成果,为推进约鲁巴语乃至非洲低资源语言的NLP研究提供基础,指明未来方向与机遇。

原文摘要 · Abstract (English)

Natural Language Processing (NLP) is becoming a dominant subset of artificial intelligence as the need to help machines understand human language looks indispensable. Several NLP applications are ubiquitous, partly due to the myriad of datasets being churned out daily through mediums like social networking sites. However, the growing development has not been evident in most African languages due to the persisting resource limitations, among other issues. Yorùbá language, a tonal and morphologically rich African language, suffers a similar fate, resulting in limited NLP usage. To encourage further research towards improving this situation, this systematic literature review aims to comprehensively analyse studies addressing NLP development for Yorùbá, identifying challenges, resources, techniques, and applications. A well-defined search string from a structured protocol was employed to search, select, and analyse 105 primary studies between 2014 and 2024 from reputable databases. The review highlights the scarcity of annotated corpora, the limited availability of pre-trained language models, and linguistic challenges like tonal complexity and diacritic dependency as significant obstacles. It also revealed the prominent techniques, including rule-based methods, among others. The findings reveal a growing body of multilingual and monolingual resources, even though the field is constrained by socio-cultural factors such as code-switching and the desertion of language for digital usage. This review synthesises existing research, providing a foundation for advancing NLP for Yorùbá and in African languages generally. It aims to guide future research by identifying gaps and opportunities, thereby contributing to the broader inclusion of Yorùbá and other under-resourced African languages in global NLP advancements.

语言模型非洲语言低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。