用21小时库尔德语语音数据训练波浪声码器,显著提升库尔德语语音合成质量。
Enhancing Kurdish Text-to-Speech with Native Corpus Training: A High-Quality WaveGlow Vocoder Approach
- 在21小时库尔德语语料上训练专用波浪声码器
- 合成语音主观评分达4.91,创库尔德语新纪录
- 为低资源语言语音合成提供可复现范式
随着文本转语音技术的进步,从文本生成语音的能力极大地促进了数字内容的可访问性。然而,对于中央库尔德语(CKB)等低资源语言而言,有效的语音合成系统开发仍面临诸多挑战,主要由于缺乏语言学信息和专用资源。本文基于Tacotron改进库尔德语语音合成系统,将波浪声码器(WaveGlow)在21小时的中央库尔德语语音语料上进行训练,而非使用预训练的英语声码器。目标语言语料上的声码器训练对于准确流畅地适应库尔德语的音素和语调变化至关重要。实验表明,该改进模型显著优于使用英语预训练模型的基线系统。特别是,自适应波浪声码器模型取得了令人印象深刻的主观评分(MOS)4.91,为库尔德语语音合成设立了新基准。本研究不仅提升了中央库尔德语的语音合成性能,也为其他库尔德语方言及相关语言的进一步发展开辟了道路。
原文摘要 · Abstract (English)
The ability to synthesize spoken language from text has greatly facilitated access to digital content with the advances in text-to-speech technology. However, effective TTS development for low-resource languages, such as Central Kurdish (CKB), still faces many challenges due mainly to the lack of linguistic information and dedicated resources. In this paper, we improve the Kurdish TTS system based on Tacotron by training the Kurdish WaveGlow vocoder on a 21-hour central Kurdish speech corpus instead of using a pre-trained English vocoder WaveGlow. Vocoder training on the target language corpus is required to accurately and fluently adapt phonetic and prosodic changes in Kurdish language. The effectiveness of these enhancements is that our model is significantly better than the baseline system with English pretrained models. In particular, our adaptive WaveGlow model achieves an impressive MOS of 4.91, which sets a new benchmark for Kurdish speech synthesis. On one hand, this study empowers the advanced features of the TTS system for Central Kurdish, and on the other hand, it opens the doors for other dialects in Kurdish and other related languages to further develop.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。