用同一模型级联生成语音增强,减少计算量且效果更好
Speech Enhancement based on cascaded two flows
- 用同一流匹配模型同时完成增强与初始值生成
- 相同或更少评估次数下性能优于或等同于基线
- 适合追求高效高质语音增强的开发者
基于扩散概率模型的语音增强表现优异,但需较高函数评估次数(NFE)。近期提出的流匹配方法在少量NFE下也具备竞争力。早期方法仅以带噪语音为条件变量;后续方法引入预测模型生成增强语音作为初始值和条件,但需额外训练预测模型。本文提出使用相同的流匹配模型,同时完成语音增强与生成初始增强语音。实验表明,该方法在采用两级生成机制的情况下,仍保持相同或更少的NFE,性能与先前基线相当或更优。
原文摘要 · Abstract (English)
Speech enhancement (SE) based on diffusion probabilistic models has exhibited impressive performance, while requiring a relatively high number of function evaluations (NFE). Recently, SE based on flow matching has been proposed, which showed competitive performance with a small NFE. Early approaches adopted the noisy speech as the only conditioning variable. There have been other approaches which utilize speech enhanced with a predictive model as another conditioning variable and to sample an initial value, but they require a separate predictive model on top of the generative SE model. In this work, we propose to employ an identical model based on flow matching for both SE and generating enhanced speech used as an initial starting point and a conditioning variable. Experimental results showed that the proposed method required the same or fewer NFEs even with two cascaded generative methods while achieving equivalent or better performances to the previous baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。