用语言模型实现情感语音合成,可精细调控情绪风格。
Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions
- 基于语言模型构建框架,通过连续维度控制情感表达。
- 在自然度和多样性上优于基线,生成更丰富的情感语音。
- 无需显式情感标签训练,适合个性化语音应用。
情感语音合成系统因情感表达的内在复杂性及现有情感标签覆盖有限,难以捕捉人类情感的全谱范围。为此,我们提出一种基于语言模型的语音合成框架,能够生成涵盖广泛情感样式的语音。该方法支持用户沿三个连续维度——愉悦度、唤醒度和支配度(PAD)——灵活控制情感表现。为实现这一目标,我们训练了一个情感维度预测器,将语音数据集中分类的情感标签映射到PAD空间,依据成熟的心理学理论。重要的是,尽管情感维度预测器利用了分类标签,但语音合成框架本身在训练过程中并不需要显式情感标签。客观与主观评估表明,该框架能有效生成更具表现力的情感样式,并在自然度和多样性方面优于基线方法。
原文摘要 · Abstract (English)
Emotional text-to-speech (TTS) systems sturggle to capture the full spectrum of human emotions due to the inherent complexity of emotional expressions and the limited coverage of existing emotion labels. To address this, we propose a language model-based TTS framework that synthesizes speech across a broad range of emotional styles. Our approach enables flexible user control along three continuous dimensions - pleasure, arousal, and dominance (PAD). To enable this, we train an emotional dimension predictor that maps categorical emotion labels in speech datasets into the PAD space, grounded in established psychological research. Importantly, while the emotional dimension predictor leverages categorical labels, the TTS framework itself does not require explict emotion labels during training. Objective and subjective evaluations demonstrate that our framework effectively generates more expressive emotional styles and enhances both naturalness and diversity compared to baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。