arXiv:2506.02280cs.AI2025-06被引 11

评估6大LLM对非洲语言支持现状,揭示98%语言被忽视

The State of Large Language Models for African Languages: Progress and Challenges

  • 对比6种大模型、8种小模型和6种专用小模型的语言覆盖情况
  • 仅42种非洲语言获支持,超98%语言未被涵盖,仅4种常被处理
  • 适合关注低资源语言、AI公平性及非洲本地化研究者阅读

大型语言模型(LLMs)正在重塑自然语言处理,但其收益对非洲2000种低资源语言几乎不存在。本文系统比较了六种大型语言模型、八种小型语言模型(SLMs)和六种专用小型语言模型(SSLMs)在非洲语言上的覆盖情况,涵盖语言支持度、训练数据集、技术限制、书写系统问题及语言建模路线图。研究识别出42种受支持的非洲语言和23个公开可用数据集。结果显示,仅有阿姆哈拉语、斯瓦希里语、南非荷兰语和马达加斯加语四种语言始终被处理,其余超过98%的非洲语言未获支持。同时,仅识别出拉丁、阿拉伯和格埃兹三种书写系统,而20种活跃书写系统被忽略。主要挑战包括数据匮乏、分词偏差、计算成本过高及评估困难。这些问题亟需推动语言标准化、社区共建语料库以及针对非洲语言的有效适配方法。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are transforming Natural Language Processing (NLP), but their benefits are largely absent for Africa's 2,000 low-resource languages. This paper comparatively analyzes African language coverage across six LLMs, eight Small Language Models (SLMs), and six Specialized SLMs (SSLMs). The evaluation covers language coverage, training sets, technical limitations, script problems, and language modelling roadmaps. The work identifies 42 supported African languages and 23 available public data sets, and it shows a big gap where four languages (Amharic, Swahili, Afrikaans, and Malagasy) are always treated while there is over 98\% of unsupported African languages. Moreover, the review shows that just Latin, Arabic, and Ge'ez scripts are identified while 20 active scripts are neglected. Some of the primary challenges are lack of data, tokenization biases, computational costs being very high, and evaluation issues. These issues demand language standardization, corpus development by the community, and effective adaptation methods for African languages.

大模型非洲语言低资源语言AI公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。