arXiv:2411.02402cs.SDcs.LG2024-11被引 5

用最优传输方法实现高质量语音转换,效果优于现有技术。

Optimal Transport Maps are Good Voice Converters

  • 基于最优传输设计多种语音表示转换算法
  • 在梅尔频谱上取得优异的FAD指标表现
  • 适用于少量参考语音数据场景,适合语音克隆应用

近期基于神经网络的最优传输映射方法已被有效应用于风格迁移。然而,其在语音转换中的应用仍不充分。本文填补这一空白,将最优传输作为语音转换框架进行研究。针对梅尔频谱和自监督语音模型的潜在表示等不同数据形式,提出多种最优传输算法。在梅尔频谱表示下,所提方法在弗雷切特音频距离(FAD)上取得显著性能提升,该结果与理论分析一致,表明该方法对目标分布与生成分布间的FAD提供了上界。在WavLM编码器的潜在空间中,即使仅有少量参考说话人数据,也达到当前最优水平,超越已有方法。

原文摘要 · Abstract (English)

Recently, neural network-based methods for computing optimal transport maps have been effectively applied to style transfer problems. However, the application of these methods to voice conversion is underexplored. In our paper, we fill this gap by investigating optimal transport as a framework for voice conversion. We present a variety of optimal transport algorithms designed for different data representations, such as mel-spectrograms and latent representation of self-supervised speech models. For the mel-spectogram data representation, we achieve strong results in terms of Frechet Audio Distance (FAD). This performance is consistent with our theoretical analysis, which suggests that our method provides an upper bound on the FAD between the target and generated distributions. Within the latent space of the WavLM encoder, we achived state-of-the-art results and outperformed existing methods even with limited reference speaker data.

语音转换最优传输自监督语音低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。