用5小时数据训练出可听懂的米佐语语音合成系统
Towards Prosodically Informed Mizo TTS without Explicit Tone Markings
- 用非自回归端到端框架,仅需5.18小时数据
- VITS模型比Tacotron2音调错误减少,主观评价更优
- 适合低资源方言语音合成研究者参考
本文报告了针对低资源、声调型的藏缅语系语言米佐语的文本转语音(TTS)系统开发。该系统仅使用5.18小时语音数据构建;在主观与客观评估中,生成语音均被认定为可感知接受且可理解。基于相同数据,分别构建了基于Tacotron2的基线模型和基于VITS的模型。在主观与客观评估中,VITS模型均优于Tacotron2模型。在音调合成方面,VITS模型表现出显著更低的音调错误率。论文表明,非自回归端到端框架可在低资源条件下实现可接受的感知质量与可理解性。
原文摘要 · Abstract (English)
This paper reports on the development of a text-to-speech (TTS) system for Mizo, a low-resource, tonal, and Tibeto-Burman language spoken primarily in the Indian state of Mizoram. The TTS was built with only 5.18 hours of data; however, in terms of subjective and objective evaluations, the outputs were considered perceptually acceptable and intelligible. A baseline model using Tacotron2 was built, and then, with the same data, another TTS model was built with VITS. In both subjective and objective evaluations, the VITS model outperformed the Tacotron2 model. In terms of tone synthesis, the VITS model showed significantly lower tone errors than the Tacotron2 model. The paper demonstrates that a non-autoregressive, end-to-end framework can achieve synthesis of acceptable perceptual quality and intelligibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。