开源100小时奥罗莫语语音数据集,助力非洲语言语音识别
Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language
- 通过众包收集100小时真实场景语音,涵盖多种口音和环境
- 用Conformer模型达15.32%的词错误率,微调Whisper后降至10.82%
- 首个公开可用的奥罗莫语语音识别数据集,适合研究者和开发者
我们提出一种面向奥罗莫语的新型自动语音识别(ASR)数据集,该语言在埃塞俄比亚及周边地区广泛使用。数据集通过众包方式收集,包含多样化的说话人和语音特征,共100小时真实环境音频及其转写文本,覆盖干净与嘈杂环境下的朗读语音。该数据集填补了奥罗莫语语音识别资源严重不足的空白。为验证其适用性,我们采用Conformer模型进行实验,混合CTC与AED损失下获得15.32%的词错误率(WER),纯CTC损失下为18.74%。此外,微调Whisper模型后,词错误率显著降低至10.82%。这些结果为奥罗莫语语音识别建立了基准,展示了提升性能的潜力与挑战。数据集已公开于https://github.com/turinaf/sagalee,欢迎用于后续研究与开发。
原文摘要 · Abstract (English)
We present a novel Automatic Speech Recognition (ASR) dataset for the Oromo language, a widely spoken language in Ethiopia and neighboring regions. The dataset was collected through a crowd-sourcing initiative, encompassing a diverse range of speakers and phonetic variations. It consists of 100 hours of real-world audio recordings paired with transcriptions, covering read speech in both clean and noisy environments. This dataset addresses the critical need for ASR resources for the Oromo language which is underrepresented. To show its applicability for the ASR task, we conducted experiments using the Conformer model, achieving a Word Error Rate (WER) of 15.32% with hybrid CTC and AED loss and WER of 18.74% with pure CTC loss. Additionally, fine-tuning the Whisper model resulted in a significantly improved WER of 10.82%. These results establish baselines for Oromo ASR, highlighting both the challenges and the potential for improving ASR performance in Oromo. The dataset is publicly available at https://github.com/turinaf/sagalee and we encourage its use for further research and development in Oromo speech processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。