构建首个大规模波斯语语音数据集,助力低资源语言语音合成。
ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages
- 设计全流程开源管道,支持波斯语语音数据采集与标注
- 生成86小时高质量音频,语音合成评分达3.76分(满分5分)
- 适合研究低资源语言语音合成与自动标注技术的学者
本研究提出ManaTTS,目前公开可用最完整的单说话人波斯语语音语料库,以及一套完整的波斯语带注释语音数据集构建框架。ManaTTS在开放的CC-0许可下发布,包含约86小时采样率为44.1 kHz的音频。同时,我们还构建了用于评估波斯语语音识别模型的VirgoolInformal数据集,时长超过5小时。该数据集配套一个完全透明、MIT许可的处理管道,包含独特的句子分词工具、有界音频分割方法及一种专为低资源语言设计的新式强制对齐技术。基于此数据集,我们训练了一个基于Tacotron2的语音合成模型,取得3.76的平均意见分数(MOS),接近使用相同声码器和自然频谱生成的语音(MOS=3.86),更接近自然波形(MOS=4.01),充分证明该语料库的高质量与有效性。
原文摘要 · Abstract (English)
In this study, we introduce ManaTTS, the most extensive publicly accessible single-speaker Persian corpus, and a comprehensive framework for collecting transcribed speech datasets for the Persian language. ManaTTS, released under the open CC-0 license, comprises approximately 86 hours of audio with a sampling rate of 44.1 kHz. Alongside ManaTTS, we also generated the VirgoolInformal dataset to evaluate Persian speech recognition models used for forced alignment, extending over 5 hours of audio. The datasets are supported by a fully transparent, MIT-licensed pipeline, a testament to innovation in the field. It includes unique tools for sentence tokenization, bounded audio segmentation, and a novel forced alignment method. This alignment technique is specifically designed for low-resource languages, addressing a crucial need in the field. With this dataset, we trained a Tacotron2-based TTS model, achieving a Mean Opinion Score (MOS) of 3.76, which is remarkably close to the MOS of 3.86 for the utterances generated by the same vocoder and natural spectrogram, and the MOS of 4.01 for the natural waveform, demonstrating the exceptional quality and effectiveness of the corpus.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。