梳理肯尼亚NLP现状,聚焦本土语言技术短板与机遇
State of NLP in Kenya: A Survey
- 系统评估肯尼亚本土语言的NLP进展与资源分布
- 指出多数土著语言在数字空间仍严重缺位
- 适合关注非洲语言科技与AI公平性的研究者
肯尼亚以其语言多样性著称,但在自然语言处理(NLP)发展上面临独特挑战与机遇,尤其体现在其未充分代表的本土语言。本综述详细评估了肯尼亚当前NLP的发展状况,重点关注基斯瓦希里语、多尔沃语、基库尤语和卢哈亚语等本地方言在数据集构建、机器翻译、情感分析和语音识别方面的进展。尽管已有部分成果,但受限于资源与工具匮乏,大多数本土语言仍缺乏数字表现。本文通过批判性评估现有数据集与NLP模型,揭示显著差距,尤其凸显大规模语言模型的缺失与本土语言数字化不足的问题。同时,分析了机器翻译、信息检索与情感分析等关键应用如何适配本地语言需求,并探讨了影响人工智能与NLP发展的治理、政策与法规,提出一项战略路线图以指导未来研究与开发。旨在为满足肯尼亚多样化语言需求的NLP技术发展奠定基础。
原文摘要 · Abstract (English)
Kenya, known for its linguistic diversity, faces unique challenges and promising opportunities in advancing Natural Language Processing (NLP) technologies, particularly for its underrepresented indigenous languages. This survey provides a detailed assessment of the current state of NLP in Kenya, emphasizing ongoing efforts in dataset creation, machine translation, sentiment analysis, and speech recognition for local dialects such as Kiswahili, Dholuo, Kikuyu, and Luhya. Despite these advancements, the development of NLP in Kenya remains constrained by limited resources and tools, resulting in the underrepresentation of most indigenous languages in digital spaces. This paper uncovers significant gaps by critically evaluating the available datasets and existing NLP models, most notably the need for large-scale language models and the insufficient digital representation of Indigenous languages. We also analyze key NLP applications: machine translation, information retrieval, and sentiment analysis-examining how they are tailored to address local linguistic needs. Furthermore, the paper explores the governance, policies, and regulations shaping the future of AI and NLP in Kenya and proposes a strategic roadmap to guide future research and development efforts. Our goal is to provide a foundation for accelerating the growth of NLP technologies that meet Kenya's diverse linguistic demands.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。