评估大模型在非洲语言上的表现,揭示数据不平等对AI影响
Where Are We? Evaluating LLM Performance on African Languages
- 构建Sahara基准,整合非洲多语言公开数据
- 多数本土语言因数据稀疏导致模型表现差
- 呼吁政策改革与包容性数据实践
非洲丰富的语言遗产在自然语言处理中长期被忽视,主要源于历史政策偏向外语,造成显著的数据不平等。本文结合非洲语言格局的理论洞察,通过Sahara——一个从大规模公开数据集中整理的综合性基准——系统评估主流大语言模型(LLMs)在非洲语言上的表现。结果表明,政策导致的数据差异直接影响模型效能;少数语言表现尚可,但多数原住民语言因数据稀疏而被边缘化。基于此,我们提出切实可行的政策与数据实践改进建议。研究强调,必须融合理论理解与实证评估,推动非洲社区AI中的语言多样性发展。
原文摘要 · Abstract (English)
Africa's rich linguistic heritage remains underrepresented in NLP, largely due to historical policies that favor foreign languages and create significant data inequities. In this paper, we integrate theoretical insights on Africa's language landscape with an empirical evaluation using Sahara - a comprehensive benchmark curated from large-scale, publicly accessible datasets capturing the continent's linguistic diversity. By systematically assessing the performance of leading large language models (LLMs) on Sahara, we demonstrate how policy-induced data variations directly impact model effectiveness across African languages. Our findings reveal that while a few languages perform reasonably well, many Indigenous languages remain marginalized due to sparse data. Leveraging these insights, we offer actionable recommendations for policy reforms and inclusive data practices. Overall, our work underscores the urgent need for a dual approach - combining theoretical understanding with empirical evaluation - to foster linguistic diversity in AI for African communities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。