用扩散变压器实现零样本语音转换,提升音色保真度与泛化能力。
Zero-shot Voice Conversion with Diffusion Transformers
- 引入外部音色扰动机制,训练时模拟真实推理场景。
- 基于参考语音上下文学习,捕捉精细音色特征,降低音色泄露。
- 支持零样本唱歌语音转换,效果媲美当前最优方法。
零样本语音转换旨在将源语音转换为未见过说话人的音色特征。传统方法存在音色泄露、音色表征不足以及训练与推理任务不匹配等问题。本文提出Seed-VC框架,通过在训练中引入外部音色扰动器,干扰源语音音色,缓解音色泄露并使训练更贴近推理场景;同时采用扩散变压器,利用完整参考语音上下文,通过上下文学习捕获细粒度音色特征。实验表明,Seed-VC优于OpenVoice和CosyVoice等强基线,在零样本语音转换任务中实现了更高的说话人相似度和更低的词错误率。进一步扩展至零样本唱歌语音转换,通过引入基频(F0)条件控制,性能达到当前最先进水平。结果验证了该方法在克服核心挑战上的有效性,为更精准、通用的语音转换系统提供新路径。
原文摘要 · Abstract (English)
Zero-shot voice conversion aims to transform a source speech utterance to match the timbre of a reference speech from an unseen speaker. Traditional approaches struggle with timbre leakage, insufficient timbre representation, and mismatches between training and inference tasks. We propose Seed-VC, a novel framework that addresses these issues by introducing an external timbre shifter during training to perturb the source speech timbre, mitigating leakage and aligning training with inference. Additionally, we employ a diffusion transformer that leverages the entire reference speech context, capturing fine-grained timbre features through in-context learning. Experiments demonstrate that Seed-VC outperforms strong baselines like OpenVoice and CosyVoice, achieving higher speaker similarity and lower word error rates in zero-shot voice conversion tasks. We further extend our approach to zero-shot singing voice conversion by incorporating fundamental frequency (F0) conditioning, resulting in comparative performance to current state-of-the-art methods. Our findings highlight the effectiveness of Seed-VC in overcoming core challenges, paving the way for more accurate and versatile voice conversion systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。