用声学特征和韵律瓶颈提升语音合成情感表达自然度
An Empirical Study on Learning Latent Representations for Emotional Speech Synthesis

- 在FastSpeech2中融合说话人嵌入与韵律瓶颈,实现情感控制
- 单说话人情感语音合成效果显著,保留身份特征
- 可迁移其他说话人风格,适合需要情感化语音的应用
近年来,深度学习推动语音合成领域快速发展,涌现出大量高可懂性和自然度的端到端语音合成系统。然而,如何有效控制语音表现力仍是挑战,生成不同风格或语调的语音成为研究热点。本文针对VLSP 2022中的情感语音合成(ESS)任务,提出解决方案:通过在FastSpeech2中引入说话人嵌入与韵律瓶颈,成功实现单一说话人的情感语音合成(子任务1),并能在仅使用目标说话人中性语音数据的情况下,将另一说话人的情感风格迁移到目标说话人身上,同时保持其身份特征(子任务2)。该方法能够生成接近真人水平、自然流畅且带有指定情感色彩的语音输出。
原文摘要 · Abstract (English)
For the last couple of years, the field of speech synthesis has improved dramatically thanks to deep learning. There are more and more deep learning-based TTS systems developed to make it possible to produce voices with high intelligibility and naturalness. Meanwhile, controlling the expressiveness is yet a big deal, generating speech in different styles or manners has received a lot of attention from community recently. This paper aims to give our solutions to deal with the task emotional speech synthesis (ESS) at VLSP 2022 which allows to generate humanlike natural-sounding voice from a given input text with desired emotional expression. By integrating speaker embedding, prosody bottleneck into FastSpeech 2, our systems can promisingly generate emotional speech of a single speaker (Sub-task 1), transfer speaking styles from another speaker to the target speaker with neutral non-expressive data while retaining the target speaker's identity (Sub-task 2).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。