用离散音高编码提升语音转换的情感表现力
Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
- 用自监督音高VQVAE提取离散音高特征,增强风格建模能力
- 在LibriTTS和ESD数据集上,音色相似度与风格迁移效果更优
- 适合需要情感化语音转换的研究者或开发者
本文提出PFlow-VC,一种基于离散音高条件的流匹配语音转换模型,通过细粒度离散音高令牌和目标说话人提示信息,实现更具表现力的语音转换。以往工作多关注说话人转换,对音色转换中的韵律与情感等表达性仍需深入探索。不同于以往方法,本研究采用简洁高效的方案:先预训练自监督音高VQVAE以离散化无关说话人的音高信息,再使用掩码音高条件流匹配模型生成梅尔频谱,使说话人转换模型具备上下文音高建模能力,显著提升语音风格迁移性能。同时,通过结合全局音色嵌入与时变音色令牌,进一步提高音色相似度。在未见的LibriTTS test-clean和情绪语音数据集ESD上的实验表明,PFlow-VC在音色转换与风格迁移方面均具优势。音频样例可在https://speechai-demo.github.io/PFlow-VC/ 查看。
原文摘要 · Abstract (English)
This paper introduces PFlow-VC, a conditional flow matching voice conversion model that leverages fine-grained discrete pitch tokens and target speaker prompt information for expressive voice conversion (VC). Previous VC works primarily focus on speaker conversion, with further exploration needed in enhancing expressiveness (such as prosody and emotion) for timbre conversion. Unlike previous methods, we adopt a simple and efficient approach to enhance the style expressiveness of voice conversion models. Specifically, we pretrain a self-supervised pitch VQVAE model to discretize speaker-irrelevant pitch information and leverage a masked pitch-conditioned flow matching model for Mel-spectrogram synthesis, which provides in-context pitch modeling capabilities for the speaker conversion model, effectively improving the voice style transfer capacity. Additionally, we improve timbre similarity by combining global timbre embeddings with time-varying timbre tokens. Experiments on unseen LibriTTS test-clean and emotional speech dataset ESD show the superiority of the PFlow-VC model in both timbre conversion and style transfer. Audio samples are available on the demo page https://speechai-demo.github.io/PFlow-VC/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。