无需训练即可实现跨语言语音转换,5秒参考音频即可达成高质量效果
Training-Free Voice Conversion with Factorized Optimal Transport
- 用分解最优传输替代kNN回归,提升特征变换效率
- 仅需5秒参考音频,在双数据集上显著提升内容保真度与鲁棒性
- 适合快速部署、资源受限场景下的语音转换应用
本文提出一种无需训练的kNN-VC改进方法Factorized MKL-VC。该方法将kNN回归替换为在WavLM嵌入空间中基于Monge-Kantorovich线性解的分解最优传输映射,解决各维度方差不均问题,实现高效特征转换。在LibriSpeech和FLEURS数据集上的实验表明,仅使用5秒参考音频,该方法在跨语言语音转换任务中显著提升内容保真度与鲁棒性,性能接近FACodec,尤其在跨语言场景表现优异。
原文摘要 · Abstract (English)
This paper introduces Factorized MKL-VC, a training-free modification for kNN-VC pipeline. In contrast with original pipeline, our algorithm performs high quality any-to-any cross-lingual voice conversion with only 5 second of reference audio. MKL-VC replaces kNN regression with a factorized optimal transport map in WavLM embedding subspaces, derived from Monge-Kantorovich Linear solution. Factorization addresses non-uniform variance across dimensions, ensuring effective feature transformation. Experiments on LibriSpeech and FLEURS datasets show MKL-VC significantly improves content preservation and robustness with short reference audio, outperforming kNN-VC. MKL-VC achieves performance comparable to FACodec, especially in cross-lingual voice conversion domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。