将语音转换推理速度提升至单步,保持高音质与音色相似性。
FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
- 通过对抗性条件扩散蒸馏,将多步扩散转为单步采样。
- 单步推理下语音质量与音色保留优于或相当传统多步方法。
- 适合需要快速生成语音的应用,如实时语音克隆。
基于扩散模型的语音转换(VC)技术如VoiceGrad因在语音质量和说话人相似性方面表现优异而受到关注。然而,其显著缺陷是反向扩散过程需数十步迭代,导致推理缓慢。为此,本文提出FastVoiceGrad,一种新型单步扩散式语音转换方法,将采样步数从数十步降至一步,同时保持多步扩散模型的高质量表现。该模型通过对抗性条件扩散蒸馏(ACDD)实现,结合生成对抗网络与扩散模型的优势,并重新设计采样初始状态。在一次性任意到任意语音转换任务上的评估表明,FastVoiceGrad在语音转换性能上优于或相当于以往多步扩散模型,同时显著提升推理速度。音频样本可访问:https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/fastvoicegrad/。
原文摘要 · Abstract (English)
Diffusion-based voice conversion (VC) techniques such as VoiceGrad have attracted interest because of their high VC performance in terms of speech quality and speaker similarity. However, a notable limitation is the slow inference caused by the multi-step reverse diffusion. Therefore, we propose FastVoiceGrad, a novel one-step diffusion-based VC that reduces the number of iterations from dozens to one while inheriting the high VC performance of the multi-step diffusion-based VC. We obtain the model using adversarial conditional diffusion distillation (ACDD), leveraging the ability of generative adversarial networks and diffusion models while reconsidering the initial states in sampling. Evaluations of one-shot any-to-any VC demonstrate that FastVoiceGrad achieves VC performance superior to or comparable to that of previous multi-step diffusion-based VC while enhancing the inference speed. Audio samples are available at https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/fastvoicegrad/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。