arXiv:2501.17612cs.SDcs.AI2025-01中稿 · ICASSP 2025被引 5

用语音提示提升零样本语音转换的相似度和自然度

VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching

  • 通过语音提示与分解特征实现上下文学习
  • 在零样本场景下提升说话人相似度与语音质量
  • 适合需要高保真语音转换的应用场景

尽管近期语音转换(VC)系统取得显著进展,但在零样本场景下提升说话人相似度仍具挑战。这源于语音中说话人特征在零样本环境下的泛化与适应困难,以及训练与推理过程间的不匹配问题。为此,我们提出VoicePrompter,一种基于语音提示的鲁棒零样本语音转换模型。该模型包含:(1) 分解语音成分的方法;(2) 基于DiT的条件流匹配(CFM)解码器,以分解特征和语音提示为条件;(3) 潜在空间混合(latent mixup),通过融合不同说话人特征增强上下文学习能力。该方法通过在潜在表示上应用混合,提升了零样本语音转换中的说话人相似度与自然度。实验表明,VoicePrompter在说话人相似度、语音可懂度和音频质量方面均优于现有零样本语音转换系统。演示地址:https://hayeong0.github.io/VoicePrompter-demo/

原文摘要 · Abstract (English)

Despite remarkable advancements in recent voice conversion (VC) systems, enhancing speaker similarity in zero-shot scenarios remains challenging. This challenge arises from the difficulty of generalizing and adapting speaker characteristics in speech within zero-shot environments, which is further complicated by mismatch between the training and inference processes. To address these challenges, we propose VoicePrompter, a robust zero-shot VC model that leverages in-context learning with voice prompts. VoicePrompter is composed of (1) a factorization method that disentangles speech components and (2) a DiT-based conditional flow matching (CFM) decoder that conditions on these factorized features and voice prompts. Additionally, (3) latent mixup is used to enhance in-context learning by combining various speaker features. This approach improves speaker similarity and naturalness in zero-shot VC by applying mixup to latent representations. Experimental results demonstrate that VoicePrompter outperforms existing zero-shot VC systems in terms of speaker similarity, speech intelligibility, and audio quality. Our demo is available at \url{https://hayeong0.github.io/VoicePrompter-demo/}.

语音转换零样本语音提示流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。