提出鲁棒奖励优化框架,提升大模型情感语音合成的自然度与真实性。
RRPO: Robust Reward Policy Optimization for LLM-based Emotional TTS
- 设计混合正则化机制,增强奖励模型对人类感知的对齐能力。
- 主观评估显示,情感表现力与自然度显著优于所有基线方法。
- 适合关注语音合成中奖励欺骗问题的研究者与开发者。
可微分强化学习框架(如 DiffRO)为可控文本转语音提供了强大支持,但在情感控制等细微任务中易受奖励欺骗影响。策略模型可能通过生成声学伪影来骗取简单奖励,从而损害听觉质量。为此,我们提出鲁棒奖励策略优化(RRPO),采用混合正则化方案,构建更贴近人类感知的鲁棒奖励模型,迫使策略放弃有害捷径,转而学习真实情感的复杂特征。消融实验证明该奖励模型具备强跨语言泛化能力。主观评价显示,该鲁棒奖励模型有效缓解了奖励欺骗,相较于所有基线,在情感表现力和自然度上均有显著提升。
原文摘要 · Abstract (English)
Differentiable reinforcement learning (RL) frameworks like DiffRO offer a powerful approach for controllable text-to-speech (TTS), but are vulnerable to reward hacking, particularly for nuanced tasks like emotion control. The policy model can exploit a vanilla Reward Model (RM) by generating acoustic artifacts to achieve spurious rewards, but at the cost of degrading perceptual quality. To address this, we propose Robust Reward Policy Optimization (RRPO), a novel framework that employs a hybrid regularization scheme. This scheme develops a robust RM whose reward signal is more reliably aligned with human perception, compelling the policy to abandon detrimental shortcuts and instead learn the complex features of genuine emotions. Our ablation study confirms the enhanced robustness of our RM, as evidenced by its strong cross-lingual generalization. The subjective evaluation demonstrates that this robust RM effectively mitigates reward hacking, leading to significant improvements in both emotional expressiveness and naturalness over all baselines. Demo page: https://lrwinr.github.io/RRPO-CosyVoice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。