用简单方法让语音合成一步生成,质量高还省训练成本
ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning
- 对预训练扩散模型逐步收紧一致性约束,实现高效单步生成
- 在LJSpeech上单步生成音质媲美顶尖方法,训练成本大幅降低
- 适合追求高效语音合成、资源受限场景的开发者使用
扩散模型在语音合成中表现优异,但通常需多步采样,推理效率低。近期研究通过将扩散模型蒸馏为一致模型,实现单步生成,但引入额外训练开销且依赖教师模型性能。本文提出ECTSpeech,首次将易一致性调优(Easy Consistency Tuning, ECT)策略应用于语音合成。通过逐步收紧预训练扩散模型的一致性约束,ECTSpeech在显著降低训练复杂度的同时,实现高质量单步生成。此外,设计多尺度门控模块(MSGate),增强去噪器在不同尺度上的特征融合能力。在LJSpeech数据集上的实验表明,ECTSpeech在单步采样下音质达到当前最优水平,同时大幅减少模型训练成本与复杂度。
原文摘要 · Abstract (English)
Diffusion models have demonstrated remarkable performance in speech synthesis, but typically require multi-step sampling, resulting in low inference efficiency. Recent studies address this issue by distilling diffusion models into consistency models, enabling efficient one-step generation. However, these approaches introduce additional training costs and rely heavily on the performance of pre-trained teacher models. In this paper, we propose ECTSpeech, a simple and effective one-step speech synthesis framework that, for the first time, incorporates the Easy Consistency Tuning (ECT) strategy into speech synthesis. By progressively tightening consistency constraints on a pre-trained diffusion model, ECTSpeech achieves high-quality one-step generation while significantly reducing training complexity. In addition, we design a multi-scale gate module (MSGate) to enhance the denoiser's ability to fuse features at different scales. Experimental results on the LJSpeech dataset demonstrate that ECTSpeech achieves audio quality comparable to state-of-the-art methods under single-step sampling, while substantially reducing the model's training cost and complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。