让语音情绪随词语动态变化,突破传统整句情绪控制局限
Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation
- 用词级情绪标注+特征调制层,实现逐词情绪控制
- 在细粒度情绪数据集上表现超越现有方法
- 适合需要精准情感表达的语音合成场景
情感文本转语音(E-TTS)是实现自然可信人机交互的核心。现有系统多依赖预设标签、参考音频或自然语言提示进行句级情绪控制,虽能表达整体情绪,却无法捕捉句内情绪动态变化。为此,我们提出Emo-FiLM,一种基于大语言模型的细粒度情绪建模框架。该框架将emotion2vec的帧级特征对齐至词语,获得词级情绪标注,并通过特征自适应线性调制(FiLM)层,直接调节文本嵌入以实现词级情绪控制。为支持评估,我们构建了细粒度情绪动态数据集FEDD,包含详尽的情绪转换标注。实验表明,Emo-FiLM在全局与细粒度任务上均优于现有方法,验证了其在表达性语音合成中的有效性与通用性。
原文摘要 · Abstract (English)
Emotional text-to-speech (E-TTS) is central to creating natural and trustworthy human-computer interaction. Existing systems typically rely on sentence-level control through predefined labels, reference audio, or natural language prompts. While effective for global emotion expression, these approaches fail to capture dynamic shifts within a sentence. To address this limitation, we introduce Emo-FiLM, a fine-grained emotion modeling framework for LLM-based TTS. Emo-FiLM aligns frame-level features from emotion2vec to words to obtain word-level emotion annotations, and maps them through a Feature-wise Linear Modulation (FiLM) layer, enabling word-level emotion control by directly modulating text embeddings. To support evaluation, we construct the Fine-grained Emotion Dynamics Dataset (FEDD) with detailed annotations of emotional transitions. Experiments show that Emo-FiLM outperforms existing approaches on both global and fine-grained tasks, demonstrating its effectiveness and generality for expressive speech synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。