构建24种非洲语言语音数据集,助力低资源语言技术发展
WAXAL: A Large-Scale Multilingual African Language Speech Corpus
- 联合非洲学术与社区组织,采集24种语言语音数据
- 包含约1250小时语音识别数据和235小时高质量语音合成数据
- 数据开源免费,适合语言保护与包容性技术研究
语音技术进步主要惠及高资源语言,导致撒哈拉以南非洲多数语言使用者面临显著数字鸿沟。为弥合这一差距,我们推出WAXAL,一个大规模、公开可获取的非洲语言语音语料库,涵盖24种语言,覆盖超过1亿听众。数据集包含两部分:约1250小时的自然语音转录文本,用于自动语音识别(ASR);约235小时由单人朗读音素均衡脚本的高质量录音,用于文本到语音(TTS)。本文详细说明了数据收集、标注与质量控制的方法,涉及与四家非洲学术及社区组织的合作。我们提供数据集的详细统计信息,并讨论其潜在局限性与伦理考量。WAXAL数据集在https://huggingface.co/datasets/google/WaxalNLP开放发布,采用宽松的CC-BY-4.0许可,旨在推动研究,促进包容性技术发展,并成为这些语言数字保存的重要资源。
原文摘要 · Abstract (English)
The advancement of speech technology has predominantly favored high-resource languages, creating a significant digital divide for speakers of most Sub-Saharan African languages. To address this gap, we introduce WAXAL, a large-scale, openly accessible speech dataset for 24 languages representing over 100 million speakers. The collection consists of two main components: an Automated Speech Recognition (ASR) dataset containing approximately 1,250 hours of transcribed, natural speech from a diverse range of speakers, and a Text-to-Speech (TTS) dataset with around 235 hours of high-quality, single-speaker recordings reading phonetically balanced scripts. This paper details our methodology for data collection, annotation, and quality control, which involved partnerships with four African academic and community organizations. We provide a detailed statistical overview of the dataset and discuss its potential limitations and ethical considerations. The WAXAL datasets are released at https://huggingface.co/datasets/google/WaxalNLP under the permissive CC-BY-4.0 license to catalyze research, enable the development of inclusive technologies, and serve as a vital resource for the digital preservation of these languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。