提出新框架,让语音合成同时表达多情绪变化和混合情绪。
Sequential Trajectories and Simultaneous Blending: Multi-Emotion Modeling for Instruction-Following TTS

- 用分组相对策略优化,结合样本自适应奖励统一训练两种多情绪任务。
- 在多情绪测试集上显著提升情绪轨迹准确性和混合情绪强度。
- 适合需要精细情感控制的语音合成应用,如虚拟角色对话。
自然语言指令可灵活控制合成语音,但现有情感语音合成系统主要只支持单个语句级情感,对多情绪控制研究不足。本文研究两类互补的多情绪语音合成任务:情绪轨迹(多个有序情感阶段)与情绪混合(多种情绪共存于整句话中)。现有方法存在监督不匹配问题:微调不显式评估情感特征,单一情感奖励无法提供轨迹完成的结构感知反馈或混合情绪的配对感知反馈。为此,本文提出 HybridEmo,一种后训练框架:先以监督微调初始化两个任务,再通过样本感知的混合奖励,使用分组相对策略优化(Group Relative Policy Optimization)对齐语音标记策略。对于轨迹样本,段落对齐一致性结合平均与最弱阶段证据,保障预设阶段的正确性与完整性;对于混合样本,基于高斯混合模型的奖励融合离线情绪空间中目标情绪锚点的帧级支持与整句级弱目标边际。两个分支共享自动语音识别(ASR)奖励,并在统一策略中路由。在 MultiEmo-Test 上,HybridEmo 显著提升情绪轨迹正确率和混合情绪强度,且说话人相似度无明显下降。人工评测显示,用户偏好优于 CosyVoice 3 与 EmoVoice-0.5B,与 Qwen3-TTS 呈近平衡偏好。
原文摘要 · Abstract (English)
Natural-language instructions enable flexible control of synthesized speech, yet emotional TTS systems primarily model a single utterance-level affect, leaving multi-emotion control underexplored. We study two complementary multi-emotion TTS tasks: emotion trajectory, which spans several ordered affective stages, and emotion blending, in which multiple emotions coexist throughout an utterance. These tasks expose a supervision mismatch: supervised fine-tuning (SFT) does not explicitly evaluate emotion features, while single-emotion rewards provide neither structure-aware feedback for trajectory completion nor pair-aware feedback for blending. We introduce HybridEmo, a post-training framework that initializes both tasks with SFT and then aligns the speech-token policy through Group Relative Policy Optimization using a sample-aware hybrid reward. For trajectory samples, segment-aligned consistency combines average and weakest-stage evidence to preserve the correctness and completeness of prescribed stages. For blending samples, a GMM-based reward combines frame-level support from the union of target-emotion anchors in an offline emotion space with an utterance-level weaker-target margin. Both branches share an ASR reward and are routed within a unified policy. On MultiEmo-Test, HybridEmo significantly improves trajectory correctness and blending intensity, without a noticeable degradation in speaker similarity. Human evaluation prefers HybridEmo to CosyVoice 3 and EmoVoice-0.5B, with nearly balanced preferences against Qwen3-TTS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。