建成首个面向6000万使用者的普什图语开源语音库,助力低资源语言语音技术发展。
Pashto Common Voice: Building the First Open Speech Corpus for a 60-Million-Speaker Low-Resource Language
- 通过社区协作与本地化界面,构建覆盖13个领域的普什图语语音数据集。
- 数据量达147小时,验证时长82.33小时,模型在测试集上词错误率降至13.4%。
- 适合低资源语言研究者、开源语音项目开发者及跨语言技术实践者使用。
我们介绍了普什图语Common Voice语料库——首个大规模、开放许可的普什图语语音资源,该语言拥有超过6000万母语使用者,此前在开源语音技术中几乎空白。2022至2025年间,通过社区努力,语料库从最初的1.5小时、5名贡献者扩展至147小时、1483位独特说话人,涵盖十次Mozilla Common Voice发布(CV14-CV23)。CV17到CV18期间,参与人数约增长108倍,与VOA普什图语广播活动同步。方法包括界面本地化、基于维基百科的句子提取与自动过滤、针对四种最常被省略的普什图语字符的音素定向采集,以及多渠道社区推广。MCV23包含107,781段语音片段(60,337段已验证,总计82.33小时验证时长),覆盖13个内容领域。在MCV20上微调Whisper Base模型,测试集词错误率为13.4%,相较公开报告的Whisper Base零样本结果99.0%有显著提升。
原文摘要 · Abstract (English)
We present the Pashto Common Voice corpus -- the first large-scale, openly licensed speech resource for Pashto, a language with over 60 million native speakers largely absent from open speech technology. Through a community effort spanning 2022-2025, the corpus grew from 1.5 hours and 5 contributors to 147 total hours and 1,483 unique speakers across ten Mozilla Common Voice releases (CV14-CV23). Speaker participation increased approximately 108-fold between CV17 and CV18, coinciding with a VOA Pashto broadcast campaign. We describe the full methodology: interface localisation, Wikipedia-based sentence extraction with automated filtering, phonemically targeted contributions for the four most frequently dropped Pashto characters, and multi-channel community outreach. MCV23 contains 107,781 clips (60,337 validated; 82.33 validated hours) across 13 content domains. Fine-tuning Whisper Base on the MCV20 yields 13.4% WER on the MCV20 test split, against the published Whisper Base zero-shot WER of 99.0% on Pashto.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。