arXiv:2509.18531eess.AScs.AI2025-09被引 3

用少量人工标注偏好优化语音韵律,让语音更自然。

No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS

  • 每轮仅需几百个真人偏好数据,直接优化语音自然度
  • 在韩语客服对话数据集上获最高人类评分,错误率也达标
  • 适合追求语音真实感的TTS系统研发者

近期研究显示,组相对策略优化(GRPO)可提升神经文本转语音(TTS)性能。然而,在缺乏可验证的语音韵律奖励时,基于转写误差(CER/NLL)训练的GRPO会将语音退化为单调、不自然的输出;加入说话人相似性后更导致训练不稳定且转写误差恶化。本文提出一种迭代式直接偏好优化(DPO)方法,仅需每轮数百个真人标注的偏好对,即可直接优化语音自然度,同时保持与当前模型的一致性。在包含真实韩语客服对话的精选数据集KoCC-TTS上,该方法在人类偏好评测(ELO)中表现最优,同时保持竞争力的转写错误率,优于GRPO和多个商用基线。结果表明,当无法自动衡量韵律质量时,人类偏好优化提供了一条高效且实用的路径,以实现自然、稳健的TTS。

原文摘要 · Abstract (English)

Recent work reports gains in neural text-to-speech (TTS) with Group Relative Policy Optimization (GRPO). However, in the absence of a verifiable reward for \textit{prosody}, GRPO trained on transcription-oriented signals (CER/NLL) lowers error rates yet collapses prosody into monotone, unnatural speech; adding speaker-similarity further destabilizes training and degrades CER. We address this with an \textit{iterative Direct Preference Optimization (DPO)} scheme that uses only a few hundred human-labeled preference pairs per round to directly optimize prosodic naturalness while regularizing to the current model. On \textbf{KoCC-TTS}, a curated dataset of authentic Korean call center interactions capturing task-oriented dialogues, our method attains the highest human preference (ELO) with competitive CER, outperforming GRPO and strong commercial baselines. These results suggest that when prosody cannot be rewarded automatically, \textit{human preference optimization} offers a practical and data-efficient path to natural and robust TTS. The demo page is available at \href{https://tts.ch.dev}

语音合成偏好优化TTS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。