50小时语音数据即可实现鲁棒的卢旺达语语音识别
How much speech data is necessary for ASR in African languages? An evaluation of data scaling in Kinyarwanda and Kikuyu
- 通过系统性数据量实验,验证小规模数据下的性能极限
- 50小时数据使错误率低于13%,200小时降至10%以下
- 发现标注噪声是主要失败原因,占高错误案例38.6%
低资源非洲语言的自动语音识别(ASR)系统开发面临转录语音数据不足的挑战。尽管近期大模型如OpenAI Whisper为低资源语言带来希望,但实际部署所需数据量及系统失效模式仍不明确。本文针对两个班图语种开展评估:在卢旺达语上进行从1到1400小时的系统性数据缩放实验;在基库尤语上使用270小时数据进行详细错误分析。结果显示,仅需50小时训练数据即可实现实用性能(词错误率WER < 13%),且性能在200小时时持续提升至WER < 10%。错误分析表明,标注质量问题是主要瓶颈,约38.6%的高错误案例源于噪声标注,凸显数据清洗与数据量同等重要。研究为类似低资源语言场景的系统构建提供可操作基准。相关模型与数据已开源。
原文摘要 · Abstract (English)
The development of Automatic Speech Recognition (ASR) systems for low-resource African languages remains challenging due to limited transcribed speech data. While recent advances in large multilingual models like OpenAI's Whisper offer promising pathways for low-resource ASR development, critical questions persist regarding practical deployment requirements. This paper addresses two fundamental concerns for practitioners: determining the minimum data volumes needed for viable performance and characterizing the primary failure modes that emerge in production systems. We evaluate Whisper's performance through comprehensive experiments on two Bantu languages: systematic data scaling analysis on Kinyarwanda using training sets from 1 to 1,400 hours, and detailed error characterization on Kikuyu using 270 hours of training data. Our scaling experiments demonstrate that practical ASR performance (WER < 13\%) becomes achievable with as little as 50 hours of training data, with substantial improvements continuing through 200 hours (WER < 10\%). Complementing these volume-focused findings, our error analysis reveals that data quality issues, particularly noisy ground truth transcriptions, account for 38.6\% of high-error cases, indicating that careful data curation is as critical as data volume for robust system performance. These results provide actionable benchmarks and deployment guidance for teams developing ASR systems across similar low-resource language contexts. We release accompanying and models see https://github.com/SunbirdAI/kinyarwanda-whisper-eval
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。