零样本语音风格转换,实现跨风格自然发音迁移。
ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training
- 用隐变量扩散模型结合语音提示,实现零样本风格迁移。
- 在4.4万小时数据上验证,生成风格多样且保真度高。
- 适合需要快速适配新语音风格的智能语音应用。
语音风格转换旨在将源语音的表达风格转换为目标风格,同时保留说话人身份。然而,以往方法多局限于情感等明确领域,限制了实际应用。本文提出ZSVC,一种新型零样本语音风格转换方法,利用语音编码器与带有语音提示机制的隐变量扩散模型,支持上下文学习实现风格转换。为解耦说话风格与说话人音色,引入信息瓶颈过滤源语音中的风格信息,并采用不确定性建模自适应实例归一化(UMAdaIN)扰动风格提示中的音色特征。此外,设计新颖的对抗训练策略,增强上下文学习能力并提升风格相似性。在44,000小时语音数据上的实验表明,ZSVC在零样本场景下能生成多样且高质量的语音风格。
原文摘要 · Abstract (English)
Style voice conversion aims to transform the speaking style of source speech into a desired style while keeping the original speaker's identity. However, previous style voice conversion approaches primarily focus on well-defined domains such as emotional aspects, limiting their practical applications. In this study, we present ZSVC, a novel Zero-shot Style Voice Conversion approach that utilizes a speech codec and a latent diffusion model with speech prompting mechanism to facilitate in-context learning for speaking style conversion. To disentangle speaking style and speaker timbre, we introduce information bottleneck to filter speaking style in the source speech and employ Uncertainty Modeling Adaptive Instance Normalization (UMAdaIN) to perturb the speaker timbre in the style prompt. Moreover, we propose a novel adversarial training strategy to enhance in-context learning and improve style similarity. Experiments conducted on 44,000 hours of speech data demonstrate the superior performance of ZSVC in generating speech with diverse speaking styles in zero-shot scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。