arXiv:2505.13805cs.SDcs.AI2025-05中稿 · InterSpeech 2025被引 8

用自然语言或参考语音控制情感强度,实现高保真情感语音转换

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech

  • 通过对比学习对齐语音与文本的情感特征
  • 融合声学特征与情感强度,提升语音自然度
  • 支持自然语言指令和参考语音双重控制,适合语音合成应用

尽管取得了显著进展,实现高保真、灵活且可解释的情感语音转换(EVC)仍具挑战。本文提出ClapFM-EVC框架,能够根据自然语言提示或参考语音生成高质量的带情感语音,并可调节情感强度。首先提出EVC-CLAP模型,基于自然语言提示和类别标签进行情感对比语言-音频预训练,以提取并对齐语音与文本中的细粒度情感元素。随后引入一个带有自适应强度门的FuEncoder,无缝融合来自预训练语音识别模型的音素后验概率(Phonetic PosteriorGrams)与情感特征。为进一步提升情感表现力与语音自然度,设计了一个条件流匹配模型,用于重建源语音的梅尔频谱图。主观与客观评估验证了该方法的有效性。

原文摘要 · Abstract (English)

Despite great advances, achieving high-fidelity emotional voice conversion (EVC) with flexible and interpretable control remains challenging. This paper introduces ClapFM-EVC, a novel EVC framework capable of generating high-quality converted speech driven by natural language prompts or reference speech with adjustable emotion intensity. We first propose EVC-CLAP, an emotional contrastive language-audio pre-training model, guided by natural language prompts and categorical labels, to extract and align fine-grained emotional elements across speech and text modalities. Then, a FuEncoder with an adaptive intensity gate is presented to seamless fuse emotional features with Phonetic PosteriorGrams from a pre-trained ASR model. To further improve emotion expressiveness and speech naturalness, we propose a flow matching model conditioned on these captured features to reconstruct Mel-spectrogram of source speech. Subjective and objective evaluations validate the effectiveness of ClapFM-EVC.

语音转换情感合成多模态可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。