用条件变分自编码器实现语音转换的多语调输出
Voice Conversion with Diverse Intonation using Conditional Variational Auto-Encoder
- 引入CVAE构建可调控语调的语音转换模型
- 通过IAF使隐空间后验更复杂,生成多样语调
- 生成语音语调丰富且音质更好,适合自然对话合成
语音转换旨在合成目标说话人的语音,同时保留源语音的语言信息。传统模型对每个输入仅生成单一结果,难以体现说话人同一文本下不同的语调变化。本文提出基于条件变分自编码器(CVAE)的新方法,将说话人风格特征映射到具有高斯分布的隐空间,并利用逆自回归流(IAF)增强隐空间后验的复杂性,从而实现同一输入生成多种语调的语音。实验表明,该方法不仅能生成多样化语调,且音质优于无CVAE的基线模型。
原文摘要 · Abstract (English)
Voice conversion is a task of synthesizing an utterance with target speaker's voice while maintaining linguistic information of the source utterance. While a speaker can produce varying utterances from a single script with different intonations, conventional voice conversion models were limited to producing only one result per source input. To overcome this limitation, we propose a novel approach for voice conversion with diverse intonations using conditional variational autoencoder (CVAE). Experiments have shown that the speaker's style feature can be mapped into a latent space with Gaussian distribution. We have also been able to convert voices with more diverse intonation by making the posterior of the latent space more complex with inverse autoregressive flow (IAF). As a result, the converted voice not only has a diversity of intonations, but also has better sound quality than the model without CVAE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。