用强化学习实现大模型语音合成的情绪精细调控
EMORL-TTS: Reinforcement Learning for Fine-Grained Emotion Control in LLM-based TTS
- 结合监督微调与强化学习,实现情绪类别、强度、强调的联合控制
- 在多个指标上优于基线,情绪准确率提升且强调更清晰
- 适合需要自然情感表达的语音合成应用,如虚拟助手、有声书
近期基于大语言模型的语音合成系统虽具备优异音质和零样本能力,但因依赖离散语音标记,难以实现精细情绪控制。现有方法或仅支持分类情绪标签,或无法适配大模型架构。我们提出EMORL-TTS,通过统一全局情感强度(VAD空间)与局部强调调节,结合监督微调与任务特定奖励驱动的强化学习,实现情绪类别、强度和强调的联合优化。进一步研究了强调位置对情绪强度的影响。实验表明,该方法在保持与强基准相当合成质量的同时,显著提升了情绪准确性、强度区分度和强调清晰度。
原文摘要 · Abstract (English)
Recent LLM-based TTS systems achieve strong quality and zero-shot ability, but lack fine-grained emotional control due to their reliance on discrete speech tokens. Existing approaches either limit emotions to categorical labels or cannot generalize to LLM-based architectures. We propose EMORL-TTS (Fine-grained Emotion-controllable TTS with Reinforcement Learning), a framework that unifies global intensity control in the VAD space with local emphasis regulation. Our method combines supervised fine-tuning with reinforcement learning guided by task-specific rewards for emotion category, intensity, and emphasis. Moreover, we further investigate how emphasis placement modulates fine-grained emotion intensity. Experiments show that EMORL-TTS improves emotion accuracy, intensity differentiation, and emphasis clarity, while preserving synthesis quality comparable to strong LLM-based baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。