用潜在扩散与流匹配模型提升非平行语音转换的音质和速度
LatentVoiceGrad: Nonparallel Voice Conversion with Latent Diffusion/Flow-Matching Models
- 在自编码器瓶颈层使用潜在扩散模型优化语音转换过程
- 流匹配模型使转换速度更快,音质保持不变
- 适合追求高音质与高速度语音转换的研究者
此前我们提出VoiceGrad,一种基于分数驱动扩散模型的非平行语音转换方法,通过训练得分网络预测不同说话人梅尔频谱的对数密度梯度,迭代调整输入频谱以逼近目标说话人。然而仍存在音质不足、转换速度较慢的问题。为此,本文将潜在扩散模型引入VoiceGrad,提出在自编码器瓶颈层进行反向扩散的新方法;同时提出以流匹配模型替代扩散模型,进一步提升转换速度而不损失质量。实验表明,新方法在语音质量与转换效率上均优于原版。
原文摘要 · Abstract (English)
Previously, we introduced VoiceGrad, a nonparallel voice conversion (VC) technique enabling mel-spectrogram conversion from source to target speakers using a score-based diffusion model. The concept involves training a score network to predict the gradient of the log density of mel-spectrograms from various speakers. VC is executed by iteratively adjusting an input mel-spectrogram until resembling the target speaker's. However, challenges persist: audio quality needs improvement, and conversion is slower compared to modern VC methods designed to operate at very high speeds. To address these, we introduce latent diffusion models into VoiceGrad, proposing an improved version with reverse diffusion in the autoencoder bottleneck. Additionally, we propose using a flow matching model as an alternative to the diffusion model to further speed up the conversion process without compromising the conversion quality. Experimental results show enhanced speech quality and accelerated conversion compared to the original.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。