用多维偏好优化提升语音合成的自然度与一致性
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
- 构建多维度偏好数据集,实现对语音质量的精细化调控
- 在多个指标上显著优于基线,智能度、发音相似度和语调均有提升
- 适合追求高保真语音生成的开发者与研究者
近年来,基于大规模语言模型的文本到语音(TTS)系统取得了显著进展,已达到接近人类水平的语音质量。融入人类反馈已被证明能有效提升系统鲁棒性。然而,现有方法在多维度偏好数据优化方面面临挑战,且常因奖励函数过自信导致性能下降。为此,我们提出多维度偏好优化(MPO),通过构建偏好集简化多维度优化数据的构建过程,实现对多个维度的精准对齐。同时,在训练中引入正则化机制,缓解基于DPO方法常见的性能退化问题。实验表明,MPO在可懂度、说话人相似度和韵律表现等方面均显著优于基线系统。
原文摘要 · Abstract (English)
In recent years, text-to-speech (TTS) has seen impressive advancements through large-scale language models, achieving human-level speech quality. Integrating human feedback has proven effective for enhancing robustness in these systems. However, current approaches face challenges in optimizing TTS with preference data across multiple dimensions and often suffer from performance degradation due to overconfidence in rewards. We propose Multidimensional Preference Optimization (MPO) to better align TTS systems with human preferences. MPO introduces a preference set that streamlines the construction of data for multidimensional preference optimization, enabling alignment with multiple dimensions. Additionally, we incorporate regularization during training to address the typical degradation issues in DPO-based approaches. Our experiments demonstrate MPO's effectiveness, showing significant improvements in intelligibility, speaker similarity, and prosody compared to baseline systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。