arXiv:2409.12403cs.CLcs.AI2024-09被引 33

用偏好对齐让语言模型语音合成更自然、更像真人。

Preference Alignment Improves Language Model-Based TTS

  • 用直接偏好优化(DPO)调整语音生成模型,使其更符合人类偏好。
  • 生成语音在清晰度、发音人相似度上超越原模型,部分指标超真人。
  • 适用于资源少的场景,且能推广到未见过的语音任务中。

近期文本转语音(TTS)技术发展表明,基于语言模型(LM)的系统已具备与传统方法相当的性能。通过偏好对齐算法进一步优化,可使语言模型更契合奖励模型所体现的人类偏好,从而提升生成内容的可接受性。本研究对偏好对齐算法(尤其是直接偏好优化,DPO)在基于语言模型的TTS中的效果进行了全面实证评估。采用参数量为11.5亿的LM-TTS模型,实验表明偏好对齐持续提升了语音的可懂度、说话人相似度以及代理主观评价得分;其中后两项指标在某些评估中甚至超过真人语音表现。研究还证明该方法适用于低资源场景,并能有效泛化至域外应用。

原文摘要 · Abstract (English)

Recent advancements in text-to-speech (TTS) have shown that language model (LM)-based systems offer competitive performance to their counterparts. Further optimization can be achieved through preference alignment algorithms, which adjust LMs to align with the preferences of reward models, enhancing the desirability of the generated content. This study presents a thorough empirical evaluation of how preference alignment algorithms, particularly Direct Preference Optimization (DPO), enhance LM-based TTS. With a 1.15B parameter LM-based TTS model, we demonstrate that preference alignment consistently improves intelligibility, speaker similarity, and proxy subjective evaluation scores, with the latter two metrics surpassing even human speech in certain evaluations. We also show preference alignment is applicable to low-resource scenarios and effectively generalized to out-of-domain applications.

语音合成偏好对齐语言模型DPO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。