分离语言与说话人特征,实现多语言自然语音合成
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
- 将语音生成拆分为语言和说话人两个独立模块
- 在多个评估指标上显著超越现有方法
- 适合需要统一说话人音色的多语言语音系统
本研究旨在实现多语言自然语音合成并保持相同说话人身份,即跨语言语音合成。其核心挑战是语言与说话人特征纠缠问题,导致跨语言系统质量落后于同语言系统。本文提出CrossSpeech++,通过有效解耦语言与说话人信息,显著提升跨语言语音合成质量。为此,我们将复杂语音生成流程分解为两个简单组件:语言相关生成器和说话人相关生成器。语言相关生成器产生不受特定说话人属性影响的语言变体;说话人相关生成器建模体现说话人身份的声学差异。通过在独立模块中处理各类信息,该方法可有效解耦语言与说话人表征。我们采用多种指标进行广泛实验,结果表明CrossSpeech++在跨语言语音合成上取得显著改进,大幅超越现有方法。
原文摘要 · Abstract (English)
The goal of this work is to generate natural speech in multiple languages while maintaining the same speaker identity, a task known as cross-lingual speech synthesis. A key challenge of cross-lingual speech synthesis is the language-speaker entanglement problem, which causes the quality of cross-lingual systems to lag behind that of intra-lingual systems. In this paper, we propose CrossSpeech++, which effectively disentangles language and speaker information and significantly improves the quality of cross-lingual speech synthesis. To this end, we break the complex speech generation pipeline into two simple components: language-dependent and speaker-dependent generators. The language-dependent generator produces linguistic variations that are not biased by specific speaker attributes. The speaker-dependent generator models acoustic variations that characterize speaker identity. By handling each type of information in separate modules, our method can effectively disentangle language and speaker representation. We conduct extensive experiments using various metrics, and demonstrate that CrossSpeech++ achieves significant improvements in cross-lingual speech synthesis, outperforming existing methods by a large margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。