统一离散与连续情绪,实现可解释的情感语音合成
UDDETTS: Unifying Discrete and Dimensional Emotions for Controllable Emotional Text-to-Speech
- 融合离散标签与三维情绪空间(ADV)进行情感控制
- 支持线性情绪调节,生成更自然的细粒度情感语音
- 适用于需要精准情感调控的研究与应用
近期大语言模型在文本到语音(TTS)领域取得显著进展,但在可控、可解释地合成细粒度情感语音方面仍面临挑战。传统方法依赖离散情绪标签控制情感类别和强度,难以捕捉人类情感感知与表达的复杂性与连续性。缺乏标注均衡且细粒度的情感语音数据集,常导致合成模型过拟合并影响情感控制效果。为此,我们提出UDDETTS——一个统一离散与维度情绪的通用大模型框架,用于可控情感语音合成。该模型引入可解释的唤醒-主导-效价(Arousal-Dominance-Valence, ADV)空间描述维度情绪,并支持基于离散情绪标签或非线性量化ADV值的情感控制。此外,设计半监督训练策略,充分整合具有不同情绪标注类型的数据集进行训练。实验表明,UDDETTS可在三个可解释维度上实现线性情感控制,具备优越的端到端情感语音合成能力。代码与演示见:https://anonymous.4open.science/w/UDDETTS。
原文摘要 · Abstract (English)
Recent large language models (LLMs) have made great progress in the field of text-to-speech (TTS), but they still face major challenges in synthesizing fine-grained emotional speech in an interpretable manner. Traditional methods rely on discrete emotion labels to control emotion categories and intensities, which cannot capture the complexity and continuity of human emotional perception and expression. The lack of large-scale emotional speech datasets with balanced emotion distributions and fine-grained emotional annotations often causes overfitting in synthesis models and impedes effective emotion control. To address these issues, we propose UDDETTS, a universal LLM framework unifying discrete and dimensional emotions for controllable emotional TTS. This model introduces the interpretable Arousal-Dominance-Valence (ADV) space for dimensional emotion description and supports emotion control driven by either discrete emotion labels or nonlinearly quantified ADV values. Furthermore, a semi-supervised training strategy is designed to comprehensively utilize diverse speech datasets with different types of emotional annotations to train the UDDETTS. Experiments show that UDDETTS achieves linear emotion control along three interpretable dimensions, and exhibits superior end-to-end emotional speech synthesis capabilities. Code and demos are available at: https://anonymous.4open.science/w/UDDETTS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。