arXiv:2412.12498cs.SDeess.AS2024-12中稿 · IEEE Transactions …被引 9

实现语音合成中从音素到语句的多层级情感精细调控。

Hierarchical Control of Emotion Rendering in Speech Synthesis

  • 基于流匹配构建情感建模框架,分层提取情感强度嵌入。
  • 在音素、词、语句三级实现可量化的情感强度控制。
  • 适合需要精确情感表达的语音合成应用,如虚拟助手、有声书。

情感文本转语音(TTS)旨在从输入文本生成逼真的情感语音。然而,多层级情感渲染的定量控制仍具挑战。本文提出一种基于流匹配的情感TTS框架,引入新颖的情感强度建模方法,实现对音素、词和语句层级情感渲染的细粒度控制。我们设计了分层情感分布(ED)提取器,可在不同语音片段层级捕获可量化的ED嵌入。同时,探索多种声学特征并评估其对情感强度建模的影响。训练过程中,分层ED嵌入有效捕捉参考音频中的情感强度变化,并与语言及说话人信息相关联。推理时,该模型不仅能生成情感语音,还能对语音成分的情感渲染进行定量控制。客观与主观评估均证明本框架在语音质量、情感表现力及分层情感控制方面具有显著效果。

原文摘要 · Abstract (English)

Emotional text-to-speech synthesis (TTS) aims to generate realistic emotional speech from input text. However, quantitatively controlling multi-level emotion rendering remains challenging. In this paper, we propose a flow-matching based emotional TTS framework with a novel approach for emotion intensity modeling to facilitate fine-grained control over emotion rendering at the phoneme, word, and utterance levels. We introduce a hierarchical emotion distribution (ED) extractor that captures a quantifiable ED embedding across different speech segment levels. Additionally, we explore various acoustic features and assess their impact on emotion intensity modeling. During TTS training, the hierarchical ED embedding effectively captures the variance in emotion intensity from the reference audio and correlates it with linguistic and speaker information. The TTS model not only generates emotional speech during inference, but also quantitatively controls the emotion rendering over the speech constituents. Both objective and subjective evaluations demonstrate the effectiveness of our framework in terms of speech quality, emotional expressiveness, and hierarchical emotion control.

语音合成情感控制流匹配多层级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。