用在线强化学习优化流匹配语音合成,提升音色相似度与听感质量
FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech

- 将微分方程路径转为随机路径,直接微调开源流模型
- 加权奖励组合更快收敛,训练去引导加速效果显著
- 适合追求音色还原与语音自然度的语音合成研究者
现有文本到语音(TTS)的强化学习研究集中于大语言模型,而对流匹配(Flow-Matching, FM)关注不足。本文提出 FlowTTS-GRPO,一种面向 FM 基础语音合成的在线强化学习框架。通过将常微分方程(ODE)轨迹转换为随机微分方程(SDE)路径,实现对开源 FM 模型的直接微调,无需辅助模型。实验表明,加权奖励组合比概率方案收敛更快;三项实用优化有效提升性能:训练中省略分类器无关引导(CFG)可加速收敛;合成困难样本增强鲁棒性;对 FM 组件应用强化学习可提升音频细节指标。在 CosyVoice 3.0 与 F5-TTS 上的实验显示,音色相似度与感知质量均获客观和主观提升,F5-TTS 还改善了语音可懂度。
原文摘要 · Abstract (English)
Existing Reinforcement Learning (RL) research for Text-to-Speech (TTS) focuses on large language models (LLMs), leaving Flow-Matching (FM) under-explored. We present FlowTTS-GRPO, an online RL framework for FM-based TTS. By converting ordinary differential equation (ODE) trajectories into stochastic differential equation (SDE) paths, our method enables direct fine-tuning of open-source FM models without auxiliary models. We show that a weighted reward combination converges faster than a probabilistic scheme, and identify three practical optimizations: omitting classifier-free guidance (CFG) during training accelerates convergence; synthesizing hard cases improves robustness; and applying RL to the FM component enhances audio-detail metrics. Experiments on CosyVoice 3.0 and F5-TTS demonstrate objective and subjective preference gains in speaker similarity and perceptual quality, with F5-TTS also improving intelligibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。