arXiv:2504.08524eess.AScs.AI2025-04

通过语义映射块消除语音转换中的音色泄漏,提升目标音色相似度。

USM-VC: Mitigating Timbre Leakage with Universal Semantic Mapping Residual Block for Voice Conversion

  • 设计通用语义映射残差块,分离内容与音色信息。
  • 用多说话人统计的音素字典生成无音色内容表征,降低泄漏。
  • 适用于多种语音转换框架,尤其适合追求音色保真的场景。

语音转换(VC)在保留语音内容的同时将源语音转换为目标音色,但源说话人的音色信息常嵌入内容表征中,导致显著的音色泄漏,降低与目标说话人的相似度。为此,我们向内容提取器引入通用语义匹配(USM)残差块,其包含两个加权分支:1)基于通用语义字典的内容特征重表达(CFR)模块,生成无音色内容表征;2)原始内容层的跳接连接,提供细粒度互补信息。在CFR模块中,通用语义字典的每个条目代表一个音素类别,通过多说话人语音统计获得,形成稳定且说话人无关的语义集合。我们提出一种CFR方法,利用对应音素后验概率作为权重,将每个内容帧表示为字典条目的加权线性组合,从而获得无音色内容表征。在多种语音转换框架上的大量实验表明,该方法能有效缓解音色泄漏,显著提升与目标说话人的音色相似度。

原文摘要 · Abstract (English)

Voice conversion (VC) transforms source speech into a target voice by preserving the content. However, timbre information from the source speaker is inherently embedded in the content representations, causing significant timbre leakage and reducing similarity to the target speaker. To address this, we introduce a Universal Semantic Matching (USM) residual block to a content extractor. The residual block consists of two weighted branches: 1) universal semantic dictionary based Content Feature Re-expression (CFR) module, supplying timbre-free content representation. 2) skip connection to the original content layer, providing complementary fine-grained information. In the CFR module, each dictionary entry in the universal semantic dictionary represents a phoneme class, computed statistically using speech from multiple speakers, creating a stable, speaker-independent semantic set. We introduce a CFR method to obtain timbre-free content representations by expressing each content frame as a weighted linear combination of dictionary entries using corresponding phoneme posteriors as weights. Extensive experiments across various VC frameworks demonstrate that our approach effectively mitigates timbre leakage and significantly improves similarity to the target speaker.

语音转换音色分离语义映射特征重表达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。