构建65小时多语言语音数据集,提升低资源语言识别准确率
WorldSpeech: A Multilingual Speech Corpus from Around the World
- 从议会记录等公开源收集76种语言语音-文本对数据
- 24种语言超1000小时,37种超200小时,平均错误率降63.5%
- 适合低资源语言ASR研究与模型微调
自动语音识别(ASR)在高资源语言上表现良好,但多数语言因缺乏公开对齐数据导致性能急剧下降。为此,我们推出了WorldSpeech,一个包含65,000小时跨76种语言的24 kHz多语言语音语料库,数据来源于议会记录、国际广播及公共领域有声书等多元公开渠道。其中37种语言拥有超过200小时对齐语音,28种超过500小时,24种超过1000小时。将现有ASR模型在WorldSpeech上微调后,在11种语言类型多样性的语言上实现平均63.5%的相对词错误率降低。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) performs well for high-resource languages with abundant paired audio-transcript data, but its accuracy degrades sharply for most languages due to limited publicly available aligned data. To this end, we introduce WorldSpeech, a 24 kHz multilingual speech corpus comprising 65k hours of aligned audio-transcript data across 76 languages, collected from diverse public sources including parliamentary proceedings, international broadcasts, and public-domain audiobooks. For 37 languages, WorldSpeech provides more than 200 hours of aligned speech, with 28 exceeding 500 hours and 24 surpassing 1k hours. Fine-tuning existing ASR models on WorldSpeech results in an average relative Word-Error-Rate reduction of 63.5% across 11 typologically diverse languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。