用直接路径生成语音,零样本转换效果好
ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization
- 基于修正流的微分方程模型,直接映射语音分布
- 小数据和零样本场景下音质高保真,性能优异
- 融合内容与音高信息优化说话人特征,更精准
近年来,基于扩散的生成模型在语音转换中表现出色,如去噪扩散概率模型(DDPM)等。然而,这些模型需大量采样步骤,限制了其在真实场景中的应用。本文提出ReFlow-VC,一种基于修正流的高质量语音转换方法。该方法为常微分方程(ODE)模型,沿最直接路径将高斯分布转换为真实的梅尔频谱分布。此外,我们提出一种新方法,通过结合内容与音高信息优化说话人特征,使特征更能反映当前语音特性。实验表明,ReFlow-VC在小数据集和零样本场景下表现卓越。
原文摘要 · Abstract (English)
In recent years, diffusion-based generative models have demonstrated remarkable performance in speech conversion, including Denoising Diffusion Probabilistic Models (DDPM) and others. However, the advantages of these models come at the cost of requiring a large number of sampling steps. This limitation hinders their practical application in real-world scenarios. In this paper, we introduce ReFlow-VC, a novel high-fidelity speech conversion method based on rectified flow. Specifically, ReFlow-VC is an Ordinary Differential Equation (ODE) model that transforms a Gaussian distribution to the true Mel-spectrogram distribution along the most direct path. Furthermore, we propose a modeling approach that optimizes speaker features by utilizing both content and pitch information, allowing speaker features to reflect the properties of the current speech more accurately. Experimental results show that ReFlow-VC performs exceptionally well in small datasets and zero-shot scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。