arXiv:2508.17868cs.SDcs.AI2025-08中稿 · Interspeech 2025被引 1

用对抗性扩散蒸馏加速语音转换,一步完成且更高效

FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation

  • 通过对抗性扩散蒸馏同时压缩扩散模型和内容编码器
  • 在单步采样下实现与快模型相当的音质与音色相似度
  • 推理速度比同类模型快6.6~6.9倍(GPU)和1.8倍(CPU)

基于扩散模型的语音转换(VC)方法(如VoiceGrad)可实现高保真语音和强说话人相似性;但其转换过程因迭代采样而缓慢。FastVoiceGrad通过将VoiceGrad蒸馏为单步扩散模型克服了该问题。然而,它仍需计算量大的内容编码器来分离说话人身份与语音内容,影响效率。为此,我们提出FasterVoiceGrad,一种通过对抗性扩散转换蒸馏(ADCD)同时蒸馏扩散模型与内容编码器的新型单步扩散式语音转换模型,蒸馏过程在转换中进行,并结合对抗训练与得分蒸馏。单次语音转换实验表明,FasterVoiceGrad在性能上与FastVoiceGrad相当,且在GPU上速度提升6.6–6.9倍,在CPU上提升1.8倍。

原文摘要 · Abstract (English)

A diffusion-based voice conversion (VC) model (e.g., VoiceGrad) can achieve high speech quality and speaker similarity; however, its conversion process is slow owing to iterative sampling. FastVoiceGrad overcomes this limitation by distilling VoiceGrad into a one-step diffusion model. However, it still requires a computationally intensive content encoder to disentangle the speaker's identity and content, which slows conversion. Therefore, we propose FasterVoiceGrad, a novel one-step diffusion-based VC model obtained by simultaneously distilling a diffusion model and content encoder using adversarial diffusion conversion distillation (ADCD), where distillation is performed in the conversion process while leveraging adversarial and score distillation training. Experimental evaluations of one-shot VC demonstrated that FasterVoiceGrad achieves competitive VC performance compared to FastVoiceGrad, with 6.6-6.9 and 1.8 times faster speed on a GPU and CPU, respectively.

语音转换扩散模型加速推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。