arXiv:2411.02026cs.SDcs.AI2024-11被引 5

通过声音特质融合与条件流匹配,实现高保真零样本语音转换

Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching

  • 分离语音内容与音色特征,用跨注意力融合多源音色嵌入
  • 在多个指标上超越现有最先进方法,提升语音自然度与相似度
  • 适合语音合成、语音克隆等需零样本迁移的应用场景

尽管零样本语音转换(VC)技术取得进展,但实现与真实录音相当的说话人相似性和自然度仍是重大挑战。本文提出CTEFM-VC框架,结合内容感知音色集成建模与条件流匹配。该方法将语音分离为内容与音色表征,并利用条件流匹配模型重建源语音的梅尔频谱图。为增强音色建模能力与生成语音自然度,我们引入上下文感知的音色集成建模方法,自适应融合多种说话人验证嵌入,并通过交叉注意力模块有效利用源内容与目标音色信息。此外,设计基于结构相似性的音色损失,实现端到端联合训练。实验表明,CTEFM-VC在所有评估指标上均持续领先,显著优于当前最先进的零样本语音转换系统。

原文摘要 · Abstract (English)

Despite recent advances in zero-shot voice conversion (VC), achieving speaker similarity and naturalness comparable to ground-truth recordings remains a significant challenge. In this letter, we propose CTEFM-VC, a zero-shot VC framework that integrates content-aware timbre ensemble modeling with conditional flow matching. Specifically, CTEFM-VC decouples utterances into content and timbre representations and leverages a conditional flow matching model to reconstruct the Mel-spectrogram of the source speech. To enhance its timbre modeling capability and naturalness of generated speech, we first introduce a context-aware timbre ensemble modeling approach that adaptively integrates diverse speaker verification embeddings and enables the effective utilization of source content and target timbre elements through a cross-attention module. Furthermore, a structural similarity-based timbre loss is presented to jointly train CTEFM-VC end-to-end. Experiments show that CTEFM-VC consistently achieves the best performance in all metrics assessing speaker similarity, speech naturalness, and intelligibility, significantly outperforming state-of-the-art zero-shot VC systems.

语音转换零样本音色建模流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。