构建了肯尼亚五种语言的3000小时语音数据集,推动非洲语言技术发展。
AfriVoices-KE: A Multilingual Speech Dataset for Kenyan Languages

- 通过手机应用收集4777名母语者语音,分脚本与即兴两部分
- 含750小时脚本语料和2250小时自然语料,覆盖11个肯尼亚相关领域
- 专为低资源语言设计,适合做语音识别与方言保护研究
AfriVoices-KE 是一个大规模多语言语音数据集,涵盖肯尼亚五种语言:Dholuo、Kikuyu、Kalenjin、Maasai 和 Somali,总时长约3,000小时。其中包含750小时脚本语音和2,250小时自发语音,由来自不同地区和背景的4,777名母语者录制。数据采集采用双路径方法:脚本语音基于整理语料库、翻译内容及特定领域生成句,覆盖11个肯尼亚相关领域;非脚本语音则通过文本与图像提示诱发,以捕捉自然语言变化和方言特征。使用定制移动应用实现手机录音。质量控制包括录音前自动信噪比验证及人工内容审核。尽管面临基础设施不稳、设备兼容性差和社区信任难题,但通过本地动员者、利益相关方合作及灵活培训方案得以缓解。该数据集为开发包容性自动语音识别与文本转语音系统提供基础资源,同时促进肯尼亚语言遗产的数字保存。
原文摘要 · Abstract (English)
AfriVoices-KE is a large-scale multilingual speech dataset comprising approximately 3,000 hours of audio across five Kenyan languages: Dholuo, Kikuyu, Kalenjin, Maasai, and Somali. The dataset includes 750 hours of scripted speech and 2,250 hours of spontaneous speech, collected from 4,777 native speakers across diverse regions and demographics. This work addresses the critical underrepresentation of African languages in speech technology by providing a high-quality, linguistically diverse resource. Data collection followed a dual methodology: scripted recordings drew from compiled text corpora, translations, and domain-specific generated sentences spanning eleven domains relevant to the Kenyan context, while unscripted speech was elicited through textual and image prompts to capture natural linguistic variation and dialectal nuances. A customized mobile application enabled contributors to record using smartphones. Quality assurance operated at multiple layers, encompassing automated signal-to-noise ratio validation prior to recording and human review for content accuracy. Though the project encountered challenges common to low-resource settings, including unreliable infrastructure, device compatibility issues, and community trust barriers, these were mitigated through local mobilizers, stakeholder partnerships, and adaptive training protocols. AfriVoices-KE provides a foundational resource for developing inclusive automatic speech recognition and text-to-speech systems, while advancing the digital preservation of Kenya's linguistic heritage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。