arXiv:2501.01861cs.SDeess.AS2025-01中稿 · 2025 IEEE Internat…被引 2

用循环一致性提升非平行语音转换的音色与音高保真度

CycleFlow: Leveraging Cycle Consistency in Flow Matching for Speaker Style Adaptation

  • 基于条件流匹配引入循环一致性约束,解决无并行数据训练难题
  • 在非平行数据上实现更自然的语音转换,音色相似度显著提升
  • 适合需要高质量音高迁移的语音合成与个性化语音应用

语音转换(VC)旨在将源说话人的音色、音高等风格特征转换为目标说话人,同时保持语言内容不变。然而,在非平行数据场景下,真实目标语音缺失导致训练-推理不一致问题。现有方法仍存在音高不准、音色适配质量低等问题,源与目标音色域间音高差异显著,易生成沙哑语音,影响转换质量。本文提出CycleFlow,一种基于条件流匹配的新型语音转换方法,通过引入循环一致性约束,在非平行数据上进行音色适配训练。进一步设计双分支流匹配模型(Dual-CFM),分别优化语音生成与音高适配。实验表明,该方法显著提升说话人相似度,生成更自然、高质量的语音。

原文摘要 · Abstract (English)

Voice Conversion (VC) aims to convert the style of a source speaker, such as timbre and pitch, to the style of any target speaker while preserving the linguistic content. However, the ground truth of the converted speech does not exist in a non-parallel VC scenario, which induces the train-inference mismatch problem. Moreover, existing methods still have an inaccurate pitch and low speaker adaptation quality, there is a significant disparity in pitch between the source and target speaker style domains. As a result, the models tend to generate speech with hoarseness, posing challenges in achieving high-quality voice conversion. In this study, we propose CycleFlow, a novel VC approach that leverages cycle consistency in conditional flow matching (CFM) for speaker timbre adaptation training on non-parallel data. Furthermore, we design a Dual-CFM based on VoiceCFM and PitchCFM to generate speech and improve speaker pitch adaptation quality. Experiments show that our method can significantly improve speaker similarity, generating natural and higher-quality speech.

语音转换流匹配音高适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。