构建1800小时非洲语言语音数据集,助力本土语音技术发展
The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages
- 采集5000+讲者语音,覆盖伊博语、豪萨语、约鲁巴语,总时长1800小时
- 在语音识别任务中平均降低75.86%的错误率,显著优于现有模型
- 专为非洲语言设计,适合研究多语言语音处理与数字包容性技术者
高质量语音技术的发展依赖大规模优质数据集。然而,包括伊博语、豪萨语和约鲁巴语在内的非洲语言仍严重缺乏数据支持,主流语音助手不支持超过2000种非洲语言,使约十亿人难以使用相关技术。尽管已有针对目标语言的数据集,但规模与多样性不足。为此,我们推出了NaijaVoices数据集,包含1800小时语音-文本数据,涵盖5000多名说话者。本文阐述了独特的数据采集方法,分析其声学多样性,并通过微调实验验证其效果:在自动语音识别任务中,平均将Whisper、MMS、XLSR模型的词错误率(WER)分别降低75.86%、52.06%、42.33%。结果表明,NaijaVoices有望推动非洲语言的多语言语音处理发展。
原文摘要 · Abstract (English)
The development of high-performing, robust, and reliable speech technologies depends on large, high-quality datasets. However, African languages -- including our focus, Igbo, Hausa, and Yoruba -- remain under-represented due to insufficient data. Popular voice-enabled technologies do not support any of the 2000+ African languages, limiting accessibility for circa one billion people. While previous dataset efforts exist for the target languages, they lack the scale and diversity needed for robust speech models. To bridge this gap, we introduce the NaijaVoices dataset, a 1,800-hour speech-text dataset with 5,000+ speakers. We outline our unique data collection approach, analyze its acoustic diversity, and demonstrate its impact through finetuning experiments on automatic speech recognition, averagely achieving 75.86% (Whisper), 52.06% (MMS), and 42.33% (XLSR) WER improvements. These results highlight NaijaVoices' potential to advance multilingual speech processing for African languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。