arXiv:2505.08278eess.AS2025-05被引 2

用自监督特征实现多语言语音转换,保持原音韵律与内容

Investigating self-supervised features for expressive, multilingual voice conversion

  • 结合自监督模型隐层特征与说话人嵌入,驱动声码器重建语音
  • 零样本转换下保留源语音内容和韵律,匹配基于音素后验图的系统性能
  • 适合多语言语音转换、无需平行数据的场景

语音转换(VC)系统广泛应用于说话人匿名化、个性化语音合成等任务。监督方法依赖并行数据学习说话人映射,但标注成本高。无监督方法通常以重建输入信号为目标,该信号包含内容与说话人信息,难以解耦,常导致说话人泄露或韵律丢失。本文探索利用自监督学习(SSL)的潜力:将多个SSL模型的隐层表示与说话人嵌入拼接后输入声码器,训练其重建原始语音。零样本语音转换结果表明,该方法能有效保留源说话人的内容与韵律特征,同时在说话人相似度上达到基于音素后验图(PPGs)的基准系统水平。

原文摘要 · Abstract (English)

Voice conversion (VC) systems are widely used for several applications, from speaker anonymisation to personalised speech synthesis. Supervised approaches learn a mapping between different speakers using parallel data, which is expensive to produce. Unsupervised approaches are typically trained to reconstruct the input signal, which is composed of the content and the speaker information. Disentangling these components is a challenge and often leads to speaker leakage or prosodic information removal. In this paper, we explore voice conversion by leveraging the potential of self-supervised learning (SSL). A combination of the latent representations of SSL models, concatenated with speaker embeddings, is fed to a vocoder which is trained to reconstruct the input. Zero-shot voice conversion results show that this approach allows to keep the prosody and content of the source speaker while matching the speaker similarity of a VC system based on phonetic posteriorgrams (PPGs).

语音转换自监督学习多语言零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。