轻量级语音合成模型,一步生成高质量语音
SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified Flow
- 基于修正流构建轻量化架构,参数大幅减少
- 仅需一步采样即达到与大模型相当的语音质量
- 适合移动端部署,兼顾效率与音质
近期基于流匹配的语音合成在提升语音质量的同时,显著减少了推理步数。本文提出SlimSpeech,一种基于修正流的轻量高效语音合成系统。在现有修正流语音合成方法基础上,优化结构以减少参数,并将其作为教师模型。通过改进重流操作,直接从大模型中导出参数更少、采样轨迹更直的小模型,并结合知识蒸馏技术进一步提升性能。实验表明,所提方法在参数显著减少的情况下,通过一步采样即可实现与大模型相当的合成效果。
原文摘要 · Abstract (English)
Recently, flow matching based speech synthesis has significantly enhanced the quality of synthesized speech while reducing the number of inference steps. In this paper, we introduce SlimSpeech, a lightweight and efficient speech synthesis system based on rectified flow. We have built upon the existing speech synthesis method utilizing the rectified flow model, modifying its structure to reduce parameters and serve as a teacher model. By refining the reflow operation, we directly derive a smaller model with a more straight sampling trajectory from the larger model, while utilizing distillation techniques to further enhance the model performance. Experimental results demonstrate that our proposed method, with significantly reduced model parameters, achieves comparable performance to larger models through one-step sampling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。