通过解耦音色与情感,实现更自然的语音情感合成。
Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement
- 用互信息引导解耦,分离参考语音中的音色与情感特征。
- 在音素级别预测情感嵌入,提升情感表达细腻度。
- 适合需要高情感真实感的语音合成场景,如虚拟角色配音。
当前的情感文本转语音(TTS)和风格迁移方法依赖参考编码器控制全局风格或情感向量,但无法捕捉参考语音的细微声学细节。为此,我们提出一种新型情感TTS方法,可在音素级别实现精细的情感嵌入预测,同时解耦参考语音的内在属性。该方法采用风格解耦机制引导两个特征提取器,降低音色与情感特征间的互信息,有效从参考语音中分离出不同风格成分。实验结果表明,该方法在生成自然且情感丰富的语音方面优于基线TTS系统。本工作凸显了解耦与细粒度表征在提升情感TTS系统质量与灵活性方面的潜力。
原文摘要 · Abstract (English)
Current emotional Text-To-Speech (TTS) and style transfer methods rely on reference encoders to control global style or emotion vectors, but do not capture nuanced acoustic details of the reference speech. To this end, we propose a novel emotional TTS method that enables fine-grained phoneme-level emotion embedding prediction while disentangling intrinsic attributes of the reference speech. The proposed method employs a style disentanglement method to guide two feature extractors, reducing mutual information between timbre and emotion features, and effectively separating distinct style components from the reference speech. Experimental results demonstrate that our method outperforms baseline TTS systems in generating natural and emotionally rich speech. This work highlights the potential of disentangled and fine-grained representations in advancing the quality and flexibility of emotional TTS systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。