arXiv:2504.02407cs.SDeess.AS2025-04被引 32

用强化学习提升流匹配语音合成的清晰度和音色相似度

F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization

  • 将流匹配语音合成转为概率分布,兼容强化学习
  • 零样本克隆下语音识别错误率降29.5%,音色相似度升4.6%
  • 适合关注语音质量与个性化合成的研究者

我们提出F5R-TTS,一种将组相对策略优化(GRPO)引入基于流匹配的文本到语音(TTS)架构的新系统。通过将流匹配TTS的确定性输出重构为概率高斯分布,该方法实现了强化学习算法的无缝集成。预训练阶段使用开源数据集训练基于F5-TTS的概率化流匹配模型;在后续强化学习阶段,采用双奖励机制:通过自动语音识别计算的词错误率(WER)和通过验证模型评估的说话人相似度(SIM)。在零样本语音克隆任务上的实验表明,与传统流匹配TTS系统相比,F5R-TTS在语音可懂度上实现29.5%的相对WER降低,在说话人相似度上实现4.6%的相对SIM分数提升。音频样例见https://frontierlabs.github.io/F5R。

原文摘要 · Abstract (English)

We present F5R-TTS, a novel text-to-speech (TTS) system that integrates Group Relative Policy Optimization (GRPO) into a flow-matching based architecture. By reformulating the deterministic outputs of flow-matching TTS into probabilistic Gaussian distributions, our approach enables seamless integration of reinforcement learning algorithms. During pretraining, we train a probabilistically reformulated flow-matching based model which is derived from F5-TTS with an open-source dataset. In the subsequent reinforcement learning (RL) phase, we employ a GRPO-driven enhancement stage that leverages dual reward metrics: word error rate (WER) computed via automatic speech recognition and speaker similarity (SIM) assessed by verification models. Experimental results on zero-shot voice cloning demonstrate that F5R-TTS achieves significant improvements in both speech intelligibility (a 29.5% relative reduction in WER) and speaker similarity (a 4.6% relative increase in SIM score) compared to conventional flow-matching based TTS systems. Audio samples are available at https://frontierlabs.github.io/F5R.

语音合成强化学习流匹配音色克隆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。