arXiv:2507.04598cs.SDeess.AS2025-07中稿 · APSIPA Transaction…

分层次预测情绪变化,实现语音合成中多粒度情感控制。

Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis

  • 分三步预测语句、词语、音素级情绪波动,利用全局上下文优化局部情感。
  • 在多个语音粒度上提升情感表现力,主观评价显示显著更自然。
  • 兼容多种语音合成系统,适合需要精细情感调控的研究与应用。

我们研究分层情绪分布(ED),以实现在文本转语音(TTS)中对情感表达的多层级定量控制。提出一种新的多步分层ED预测模块,量化语句、词语和音素级别的情绪变化。通过多步骤预测,利用全局情感上下文来细化局部情绪差异,从而捕捉语音情感的内在层次结构。该方法集成到变异数适配器及外部模块设计中,兼容多种TTS系统。客观与主观评估均表明,所提框架显著增强情感表现力,并可在多个语音粒度上实现精准的情感渲染控制。

原文摘要 · Abstract (English)

We investigate hierarchical emotion distribution (ED) for achieving multi-level quantitative control of emotion rendering in text-to-speech synthesis (TTS). We introduce a novel multi-step hierarchical ED prediction module that quantifies emotion variance at the utterance, word, and phoneme levels. By predicting emotion variance in a multi-step manner, we leverage global emotional context to refine local emotional variations, thereby capturing the intrinsic hierarchical structure of speech emotion. Our approach is validated through its integration into a variance adaptor and an external module design compatible with various TTS systems. Both objective and subjective evaluations demonstrate that the proposed framework significantly enhances emotional expressiveness and enables precise control of emotion rendering across multiple speech granularities.

语音合成情感控制分层建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。