无需文本和成对数据,实现情感风格迁移的语音转换框架。
Textless and Non-Parallel Speech-to-Speech Emotion Style Transfer
- 通过分析-合成管道提取语义、说话人和情感特征,实现零样本迁移。
- 在无文本、非平行数据下,情感迁移效果优于已有方法。
- 适用于情感识别的数据增强,提升模型鲁棒性。
给定一对源语音和参考语音,语音到语音(S2S)情感风格迁移旨在生成模仿参考语音情感特征,同时保留源语音内容和说话人属性的输出语音。本文提出一种零样本语音到语音情感风格迁移框架,称为 S2S-ZEST,可在不依赖文本和成对数据的情况下,将参考语音的情感特性迁移到源语音中,同时保持说话人身份和语音内容。该框架采用分析-合成流程:分析模块从语音中提取语义标记、说话人表征和情感嵌入,并学习音高轮廓估计器与持续时间预测器;合成模块基于输入表征与推导出的声学因子生成语音。整个流程通过自编码目标进行训练,以实现推理阶段高效重合成。在实际迁移中,使用参考语音提取的情感嵌入与源语音的其余表征共同输入合成模块,生成风格迁移后的语音。实验评估显示,该方法在内容与说话人保留度上表现良好,且在情感风格迁移有效性方面优于先前方法。此外,本工作还可用于情感识别任务中的数据增强。
原文摘要 · Abstract (English)
Given a pair of source and reference speech recordings, speech-to-speech (S2S) emotion style transfer involves the generation of an output speech that mimics the emotion characteristics of the reference while preserving the content and speaker attributes of the source. In this paper, we propose a speech-to-speech zero-shot emotion style transfer framework, termed S2S Zero-shot Emotion Style Transfer (S2S-ZEST), that enables the transfer of emotional attributes from the reference to the source while retaining the speaker identity and speech content. The S2S-ZEST framework consists of an analysis-synthesis pipeline in which the analysis module extracts semantic tokens, speaker representations, and emotion embeddings from speech. Using these representations, a pitch contour estimator and a duration predictor are learned. Further, a synthesis module is designed to generate speech based on the input representations and the derived factors. The analysis-synthesis pipeline is trained using an auto-encoding objective to enable efficient resynthesis during inference. For S2S emotion style transfer, the emotion embedding extracted from the reference speech along with the remaining representations from the source speech are used in the synthesis module to generate the style-transferred speech. In our experiments, we evaluate the converted speech on content and speaker preservation (with respect to the source) as well as on the effectiveness of the emotion style transfer (with respect to the reference). The proposed framework demonstrates improved emotion style transfer performance over prior methods in a textless and non-parallel setting. We also illustrate the application of the proposed work for data augmentation in emotion recognition tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。