用预训练模型分离语义与韵律,提升零样本语音转换保真度
Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
- 基于流匹配生成模型,结合掩码重建训练实现语义与音色解耦
- 在LibriTTS上提升说话人相似度,情绪识别模型增强源语音韵律保留
- 适合关注语音转换中音色与情感保持的研究者
语音转换(VC)旨在改变说话人音色的同时保留语音内容。先前方法将自监督模型输出分词为语义标记,促进语音内容信息的解耦。近期,上下文学习(ICL)在文本转语音系统中被用于通过上下文条件建模特定特征如音色。本文提出一种增强ICL能力的语音转换系统(ICL-VC),采用基于流匹配生成模型的掩码与重建训练策略。结合语义标记,在LibriTTS数据集上的实验表明,ICL-VC提升了说话人相似度。此外,我们发现k-means是一种适用于多种预训练模型的通用分词方法。然而,ICL-VC系统在保留源语音韵律方面存在挑战。为缓解此问题,我们提出将从预训练情绪识别模型提取的韵律嵌入引入系统。在情感语音数据库上的验证表明,该集成显著增强了系统对源语音韵律的保持能力。
原文摘要 · Abstract (English)
Voice conversion (VC) aims to modify the speaker's timbre while retaining speech content. Previous approaches have tokenized the outputs from self-supervised into semantic tokens, facilitating disentanglement of speech content information. Recently, in-context learning (ICL) has emerged in text-to-speech (TTS) systems for effectively modeling specific characteristics such as timbre through context conditioning. This paper proposes an ICL capability enhanced VC system (ICL-VC) employing a mask and reconstruction training strategy based on flow-matching generative models. Augmented with semantic tokens, our experiments on the LibriTTS dataset demonstrate that ICL-VC improves speaker similarity. Additionally, we find that k-means is a versatile tokenization method applicable to various pre-trained models. However, the ICL-VC system faces challenges in preserving the prosody of the source speech. To mitigate this issue, we propose incorporating prosody embeddings extracted from a pre-trained emotion recognition model into our system. Integration of prosody embeddings notably enhances the system's capability to preserve source speech prosody, as validated on the Emotional Speech Database.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。