无需标注数据,用球面向量控制语音情绪风格与强度。
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
- 用自适应球面向量建模情绪风格与强度,无需人工标注。
- 在零样本场景下,对新说话人和情绪风格均有良好泛化能力。
- 适合需要灵活控制语音情绪的语音合成应用,如虚拟角色配音。
近年来,情感文本转语音(TTS)技术取得显著进展,但仍面临情绪本质复杂及可用情感语音数据集与模型有限的挑战。以往研究多依赖有限的情感语音数据集或需大量人工标注,限制了其在不同说话人和情感风格间的泛化能力。本文提出 EmoSphere++,一种可控制情绪风格与强度的零样本情感 TTS 模型。引入新型情绪自适应球面向量,无需人工标注即可建模情绪风格与强度;设计多层级风格编码器,确保对已见与未见说话人的有效泛化;添加额外损失函数,提升零样本场景下的情绪迁移性能。采用基于条件流匹配的解码器,在少量采样步骤内实现高质量、富有表现力的情感语音合成。实验结果验证了该框架的有效性。
原文摘要 · Abstract (English)
Emotional text-to-speech (TTS) technology has achieved significant progress in recent years; however, challenges remain owing to the inherent complexity of emotions and limitations of the available emotional speech datasets and models. Previous studies typically relied on limited emotional speech datasets or required extensive manual annotations, restricting their ability to generalize across different speakers and emotional styles. In this paper, we present EmoSphere++, an emotion-controllable zero-shot TTS model that can control emotional style and intensity to resemble natural human speech. We introduce a novel emotion-adaptive spherical vector that models emotional style and intensity without human annotation. Moreover, we propose a multi-level style encoder that can ensure effective generalization for both seen and unseen speakers. We also introduce additional loss functions to enhance the emotion transfer performance for zero-shot scenarios. We employ a conditional flow matching-based decoder to achieve high-quality and expressive emotional TTS in a few sampling steps. Experimental results demonstrate the effectiveness of the proposed framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。