提出词级语调模型,提升多语言语音合成自然度
Word-wise intonation model for cross-language TTS systems
- 通过动态时间规整聚类简化音高,降低重读音节位置差异影响
- 支持规则与语言模型两种方式预测语调轮廓,适配多语言系统
- 模型对参数变化鲁棒,适合语音合成与语调研究
本文提出一种针对俄语的词级语调模型,并展示其在其他语言中的可推广性。该模型适用于自动数据标注,可扩展用于文本到语音系统中的语调轮廓建模,既可通过基于规则的算法实现,也可通过语言模型进行预测。核心思路是通过音高简化与动态时间规整聚类相结合,部分消除因单词中重读音节位置不同带来的语调变异性。模型可作为语调研究工具,或作为语音合成中韵律描述的基础架构。我们展示了该模型与现有语调系统的关联性,以及利用语言模型进行韵律预测的可能性。最后,通过实验验证了系统对参数变化的鲁棒性。
原文摘要 · Abstract (English)
In this paper we propose a word-wise intonation model for Russian language and show how it can be generalized for other languages. The proposed model is suitable for automatic data markup and its extended application to text-to-speech systems. It can also be implemented for an intonation contour modeling by using rule-based algorithms or by predicting contours with language models. The key idea is a partial elimination of the variability connected with different placements of a stressed syllable in a word. It is achieved with simultaneous applying of pitch simplification with a dynamic time warping clustering. The proposed model could be used as a tool for intonation research or as a backbone for prosody description in text-to-speech systems. As the advantage of the model, we show its relations with the existing intonation systems as well as the possibility of using language models for prosody prediction. Finally, we demonstrate some practical evidence of the system robustness to parameter variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。