对比4种ASR模型在13种非洲语言上的表现,指导低资源场景选型。
Benchmarking Automatic Speech Recognition Models for African Languages
- 在1-400小时数据上微调4个主流ASR模型,系统评估性能。
- MMS/W2v-BERT适合极低资源,XLS-R随数据增长表现更优,Whisper在中等资源下最佳。
- 揭示外部语言模型解码的增益与失效场景,助于优化部署策略。
非洲语言的自动语音识别(ASR)受限于标注数据稀缺及缺乏系统性的模型选择、数据扩展和解码策略指导。尽管Whisper、XLS-R、MMS和W2v-BERT等大模型提升了可访问性,其在非洲低资源环境下的表现仍未被统一系统研究。本文在13种非洲语言上基准测试4种先进ASR模型,使用从1到400小时的逐步扩大的标注数据集进行微调。除报告错误率外,还揭示模型在不同条件下的行为差异:在极低资源下MMS和W2v-BERT更具数据效率;随着数据增加,XLS-R表现更优;而Whisper在中等资源条件下优势明显。同时分析外部语言模型解码的提升效果,发现其增益取决于声学与文本资源对齐程度,某些情况下会停滞甚至引入错误。本研究揭示了预训练覆盖范围、模型结构、数据领域与资源可用性间的相互作用,为非主流语言的ASR系统设计提供实用指导。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) for African languages remains constrained by limited labeled data and the lack of systematic guidance on model selection, data scaling, and decoding strategies. Large pre-trained systems such as Whisper, XLS-R, MMS, and W2v-BERT have expanded access to ASR technology, but their comparative behavior in African low-resource contexts has not been studied in a unified and systematic way. In this work, we benchmark four state-of-the-art ASR models across 13 African languages, fine-tuning them on progressively larger subsets of transcribed data ranging from 1 to 400 hours. Beyond reporting error rates, we provide new insights into why models behave differently under varying conditions. We show that MMS and W2v-BERT are more data efficient in very low-resource regimes, XLS-R scales more effectively as additional data becomes available, and Whisper demonstrates advantages in mid-resource conditions. We also analyze where external language model decoding yields improvements and identify cases where it plateaus or introduces additional errors, depending on the alignment between acoustic and text resources. By highlighting the interaction between pre-training coverage, model architecture, dataset domain, and resource availability, this study offers practical and insights into the design of ASR systems for underrepresented languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。