系统梳理非洲低资源语言语音识别研究现状与挑战。
Automatic Speech Recognition (ASR) for African Low-Resource Languages: A Systematic Literature Review
- 按PRISMA标准筛选71篇论文,覆盖111种语言的74个数据集。
- 总语音时长约1.1万小时,但超85%研究未提供可复现材料。
- 呼吁建立伦理平衡、轻量化模型与共享基准的可持续生态。
语音识别(ASR)在全球取得显著进展,但非洲低资源语言仍严重缺乏支持,阻碍了大陆数字包容性,该地区超过2000种语言。本文采用PRISMA 2020流程,系统回顾2020年1月至2025年7月间在DBLP、ACM Digital Library、Google Scholar、Semantic Scholar和arXiv上发表的研究,聚焦非洲语言的语音数据集、模型与训练方法、评估技术、挑战及未来方向。共筛查2062条记录,最终纳入71篇研究,涵盖74个数据集,覆盖111种语言,总计约11,206小时语音。不足15%的研究提供可复现材料,数据集许可不明确。自监督与迁移学习具潜力,但受限于预训练数据少、方言覆盖不足及资源匮乏。多数研究仅使用词错误率(WER),极少采用字符错误率(CER)或声调错误率(DER)等语言学导向指标,难以有效评估声调丰富及形态复杂的语言。现有证据不一致,受制于数据可用性差、标注质量低、许可模糊及基准缺失。然而,社区驱动项目与方法创新为改善带来希望。可持续发展需推动利益相关方合作、构建伦理均衡的数据集、采用轻量化建模与主动基准测试。
原文摘要 · Abstract (English)
ASR has achieved remarkable global progress, yet African low-resource languages remain rigorously underrepresented, producing barriers to digital inclusion across the continent with more than +2000 languages. This systematic literature review (SLR) explores research on ASR for African languages with a focus on datasets, models and training methods, evaluation techniques, challenges, and recommends future directions. We employ the PRISMA 2020 procedures and search DBLP, ACM Digital Library, Google Scholar, Semantic Scholar, and arXiv for studies published between January 2020 and July 2025. We include studies related to ASR datasets, models or metrics for African languages, while excluding non-African, duplicates, and low-quality studies (score <3/5). We screen 71 out of 2,062 records and we record a total of 74 datasets across 111 languages, encompassing approximately 11,206 hours of speech. Fewer than 15% of research provided reproducible materials, and dataset licensing is not clear. Self-supervised and transfer learning techniques are promising, but are hindered by limited pre-training data, inadequate coverage of dialects, and the availability of resources. Most of the researchers use Word Error Rate (WER), with very minimal use of linguistically informed scores such as Character Error Rate (CER) or Diacritic Error Rate (DER), and thus with limited application in tonal and morphologically rich languages. The existing evidence on ASR systems is inconsistent, hindered by issues like dataset availability, poor annotations, licensing uncertainties, and limited benchmarking. Nevertheless, the rise of community-driven initiatives and methodological advancements indicates a pathway for improvement. Sustainable development for this area will also include stakeholder partnership, creation of ethically well-balanced datasets, use of lightweight modelling techniques, and active benchmarking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。