arXiv:2410.00035eess.AScs.CL2024-10中稿 · ICNLSP 2024

构建60小时乌兹别克语朗读语音库,提升语音识别准确率

FeruzaSpeech: A 60 Hour Uzbek Read Speech Corpus with Punctuation, Casing, and Context

  • 采集单名母语女性主播的60小时高质量朗读录音
  • 整合后使多个数据集的词错误率显著降低
  • 支持西里尔与拉丁字母双格式,适合语言研究者使用

本文介绍FeruzaSpeech,一个乌兹别克语朗读语音语料库,包含西里尔与拉丁字母双版本转录文本,免费用于学术研究。语料库由乌兹别克斯坦塔什干的一位母语女性主播录制,共60小时,内容来自书籍节选和BBC新闻。实验表明,将FeruzaSpeech融入CommonVoice 16.1的乌兹别克语数据、乌兹别克语语音语料库数据及自身数据后,各数据集的词错误率(WER)均显著下降。

原文摘要 · Abstract (English)

This paper introduces FeruzaSpeech, a read speech corpus of the Uzbek language, containing transcripts in both Cyrillic and Latin alphabets, freely available for academic research purposes. This corpus includes 60 hours of high-quality recordings from a single native female speaker from Tashkent, Uzbekistan. These recordings consist of short excerpts from a book and BBC News. This paper discusses the enhancement of the Word Error Rates (WERs) on CommonVoice 16.1's Uzbek data, Uzbek Speech Corpus data, and FeruzaSpeech data upon integrating FeruzaSpeech.

语音语料库乌兹别克语语音识别多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。