arXiv:2606.28249eess.AScs.CL2026-06被引 1

让语音更富情感:通过分离内容与风格,提升文本转语音的表达力。

HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech

论文配图:HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech
图 1 · 摘自论文原文
  • 分离语音内容与情感风格,避免优化冲突
  • 逐级对齐帧、词、句级别目标,解决奖励稀疏问题
  • 兼顾情感丰富性与语音可懂度,适合情感化语音生成

基于大语言模型的文本转语音(TTS)虽已实现高度自然,但标准监督微调常收敛于平均语调,限制情感表现力。现有偏好驱动优化存在双重结构缺陷:内容与情感共享隐空间导致梯度冲突,引发奖励欺骗和语义退化;句子级稀疏奖励难以指导帧级生成。为此,本文提出HPRO框架,引入新型可微分奖励模型HD-Emo Codec,将语音分解为独立的内容与风格偏好标记,结构化隔离情感优化与语义内容。在此基础上,HPRO通过逐级对齐帧、词、句级别目标,弥合尺度差距。实验表明,HPRO显著增强情感表现力,同时有效保持语言可理解性。代码与音频样例公开于https://xxh333.github.io/hpro-demo/。

原文摘要 · Abstract (English)

Recently, Large Language Model (LLM)-based Text-to-Speech (TTS) models have achieved remarkable naturalness. However, the standard Supervised Fine-Tuning paradigm often converges to statistically averaged prosody, limiting emotional expressiveness. While preference-driven optimization offers a promising alternative, existing approaches suffer from two structural mismatches: information conflict, where content and emotion in a shared latent space produce conflicting gradients, leading to reward hacking and semantic degradation; and scale gap, where sparse sentence-level rewards struggle to guide dense frame-level generation. To overcome these challenges, we propose HPRO, a hierarchical progressive reward optimization framework. Within HPRO, we introduce the HD-Emo codec as a novel differentiable reward model to resolve the information conflict. It extracts speech into distinct content and style preference tokens, structurally isolating emotional optimization from semantic content. Building upon this structured preference space, HPRO bridges the scale gap by progressively aligning frame-, word- and sentence-level objectives. Experiments demonstrate that HPRO significantly enhances emotional expressiveness, while effectively preserving linguistic intelligibility. The code and audio samples are publicly available at https://xxh333.github.io/hpro-demo/.

语音生成情感表达强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。