用编码器+扩散模型实现零样本语音转换,音色更自然清晰
CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion
- 结合语音编码器与扩散模型,端到端分离语言内容与音色
- 引入混合风格归一化与多尺度音色建模,提升音色一致性与细节还原
- 双无分类器引导生成,显著改善音色相似度与语音质量
零样本语音转换旨在将源说话人的音色转换为目标说话人,同时保持语言内容不变。现有主流方法依赖预训练识别模型分离语言内容与说话人表征,导致解耦后的语言内容中残留音色信息,且说话人表征建模不足。本文提出 CoDiff-VC,一种融合语音编码器与扩散模型的端到端零样本语音转换框架,可生成高保真波形。通过单码本编码器分离语言内容与源语音,引入混合风格归一化(MSLN)扰动原始音色以增强内容解耦;采用多尺度说话人音色建模方法,确保音色一致性并提升语音细节相似性。为提高语音质量和说话人相似度,引入双无分类器引导机制,在生成过程中同时提供内容与音色引导。客观与主观实验均表明,CoDiff-VC 显著提升说话人相似度,生成自然且高质量的语音。
原文摘要 · Abstract (English)
Zero-shot voice conversion (VC) aims to convert the original speaker's timbre to any target speaker while keeping the linguistic content. Current mainstream zero-shot voice conversion approaches depend on pre-trained recognition models to disentangle linguistic content and speaker representation. This results in a timbre residue within the decoupled linguistic content and inadequacies in speaker representation modeling. In this study, we propose CoDiff-VC, an end-to-end framework for zero-shot voice conversion that integrates a speech codec and a diffusion model to produce high-fidelity waveforms. Our approach involves employing a single-codebook codec to separate linguistic content from the source speech. To enhance content disentanglement, we introduce Mix-Style layer normalization (MSLN) to perturb the original timbre. Additionally, we incorporate a multi-scale speaker timbre modeling approach to ensure timbre consistency and improve voice detail similarity. To improve speech quality and speaker similarity, we introduce dual classifier-free guidance, providing both content and timbre guidance during the generation process. Objective and subjective experiments affirm that CoDiff-VC significantly improves speaker similarity, generating natural and higher-quality speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。