用人类反馈强化学习优化扩散语音模型,提升音质与效率。
Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback
- 将原始训练损失融入奖励函数,指导语音生成优化。
- 在WaveGrad 2上实现UTMOS 3.65、NISQA 4.02的客观指标提升。
- 适合追求高保真、低延迟语音合成的实时应用开发者。
扩散模型虽能生成高保真语音,但因去噪步骤冗长且难以建模语调与节奏,难以实现实时应用。为此,本文提出扩散损失引导策略优化(DLPO),一种面向文本到语音扩散模型的强化学习与人类反馈框架。DLPO将原始训练损失融入奖励函数,在保留生成能力的同时减少效率瓶颈。通过自然度评分作为反馈信号,使奖励优化与扩散模型结构对齐,显著提升语音质量。我们在非自回归扩散模型WaveGrad 2上评估了DLPO,结果表明其在客观指标(UTMOS 3.65,NISQA 4.02)和主观评测中均表现优异,有67%的音频被用户偏好。这些发现证明了DLPO在资源受限、实时场景下的高效高质量语音合成潜力。
原文摘要 · Abstract (English)
Diffusion models produce high-fidelity speech but are inefficient for real-time use due to long denoising steps and challenges in modeling intonation and rhythm. To improve this, we propose Diffusion Loss-Guided Policy Optimization (DLPO), an RLHF framework for TTS diffusion models. DLPO integrates the original training loss into the reward function, preserving generative capabilities while reducing inefficiencies. Using naturalness scores as feedback, DLPO aligns reward optimization with the diffusion model's structure, improving speech quality. We evaluate DLPO on WaveGrad 2, a non-autoregressive diffusion-based TTS model. Results show significant improvements in objective metrics (UTMOS 3.65, NISQA 4.02) and subjective evaluations, with DLPO audio preferred 67\% of the time. These findings demonstrate DLPO's potential for efficient, high-quality diffusion TTS in real-time, resource-limited settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。