用强化学习提升语音编辑与零样本语音合成能力
CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS

- 两阶段训练:先监督微调,再无目标语音的强化优化
- 在多个评测中同时提升语音编辑和零样本语音合成效果
- 适合语音生成、语音编辑方向研究者参考
语音编辑与零样本文本转语音(TTS)共享基于语音提示的生成基础,但语音编辑对局部声学一致性要求更高。以往方法依赖监督微调(SFT),受限于不完美的成对编辑数据和粗粒度优化信号。为此,我们提出CosyEdit2,采用两阶段后训练框架:从监督编辑初始化出发,再在无目标语音数据上进行面向编辑的组相对策略优化(GRPO)。大量实验表明,CosyEdit2不仅显著提升语音编辑性能,还解锁了更优的零样本TTS能力,揭示了两项任务间的深层协同关系。音频样例见 https://cjy1018.github.io/CosyEdit2。
原文摘要 · Abstract (English)
Speech editing and zero-shot Text-to-Speech (TTS) share a similar generative foundation conditioned on speech prompts, yet speech editing demands far stricter local acoustic consistency with surrounding unedited content. While prior work has shown that Supervised Fine-Tuning (SFT) enables TTS models to acquire functional editing capability, this approach remains fundamentally bottlenecked by imperfect paired editing data and coarse-grained optimization signals. To address these limitations, we propose CosyEdit2, a speech editing model built on a two-stage post-training framework that progresses from supervised editing initialization to editing-oriented Group Relative Policy Optimization (GRPO) over target-speech-free data. Extensive experiments demonstrate that CosyEdit2 not only substantially advances speech editing performance, but also unlocks better zero-shot TTS capability, revealing a deeper mutual relationship between the two tasks. Audio samples are available at https://cjy1018.github.io/CosyEdit2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。