arXiv:2409.01995eess.AScs.AI2024-09被引 12

用离散语音符号直接生成语音,实现高质量任意说话人转换。

vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders

  • 用自监督模型的离散符号做语音内容,参考提示控制音色变换。
  • 无需标注数据训练,音质和音色相似度显著优于现有方法。
  • 支持跨语言转换,适合语音合成与个性化语音克隆场景。

我们提出一种新的语音离散符号声码器 vec2wav 2.0,推动语音转换(VC)发展。采用自监督模型生成的离散符号作为源语音内容特征,并将 VC 视为一种提示驱动的声码任务。为弥补内容符号中缺失的说话人音色信息,vec2wav 2.0 引入 WavLM 特征提供强音色依赖信息。提出新型自适应 Snake 激活函数,更有效地将音色融入波形重建过程。该方法能根据不同参考提示适配性地改变音色。整个模型无需任何监督数据即可有效训练。实验表明,在任意说话人转换任务中,vec2wav 2.0 在音频质量和说话人相似度上均显著超越所有基线。消融实验证实了所提技术的有效性。此外,仅在单语语料库上训练即可实现具有竞争力的跨语言语音转换。结果表明,仅通过语音符号声码器即可操控音色,推动语音转换与语音合成的边界。

原文摘要 · Abstract (English)

We propose a new speech discrete token vocoder, vec2wav 2.0, which advances voice conversion (VC). We use discrete tokens from speech self-supervised models as the content features of source speech, and treat VC as a prompted vocoding task. To amend the loss of speaker timbre in the content tokens, vec2wav 2.0 utilizes the WavLM features to provide strong timbre-dependent information. A novel adaptive Snake activation function is proposed to better incorporate timbre into the waveform reconstruction process. In this way, vec2wav 2.0 learns to alter the speaker timbre appropriately given different reference prompts. Also, no supervised data is required for vec2wav 2.0 to be effectively trained. Experimental results demonstrate that vec2wav 2.0 outperforms all other baselines to a considerable margin in terms of audio quality and speaker similarity in any-to-any VC. Ablation studies verify the effects made by the proposed techniques. Moreover, vec2wav 2.0 achieves competitive cross-lingual VC even only trained on monolingual corpus. Thus, vec2wav 2.0 shows timbre can potentially be manipulated only by speech token vocoders, pushing the frontiers of VC and speech synthesis.

语音转换声码器自监督音色控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。