用一致性蒸馏实现1步语音增强,速度提升54倍且效果更优
Robust One-step Speech Enhancement via Consistency Distillation
- 引入随机学习轨迹提升抗噪能力
- 结合时域辅助损失超越教师模型性能
- 首个纯一步一致性蒸馏语音增强模型
扩散模型在语音增强中表现优异,但多步迭代采样限制了其实时应用。一致性蒸馏通过从多步扩散教师模型中蒸馏出一步一致性模型成为新方向。然而,蒸馏模型天然偏向教师模型的采样轨迹,对噪声鲁棒性差且易继承教师误差。为此,本文提出ROSE-CD:一种鲁棒的一步一致性蒸馏语音增强方法。通过引入随机学习轨迹提升抗噪能力,并联合优化时域辅助损失,使模型能纠正教师引入的误差。这是首个纯一步一致性蒸馏的扩散语音增强模型,在VoiceBank-DEMAND数据集上达到当前最优语音质量,推理速度比30步教师模型快54倍。在域外数据集和真实环境噪声录音上也验证了其泛化能力。
原文摘要 · Abstract (English)
Diffusion models have shown strong performance in speech enhancement, but their real-time applicability has been limited by multi-step iterative sampling. Consistency distillation has recently emerged as a promising alternative by distilling a one-step consistency model from a multi-step diffusion-based teacher model. However, distilled consistency models are inherently biased towards the sampling trajectory of the teacher model, making them less robust to noise and prone to inheriting inaccuracies from the teacher model. To address this limitation, we propose ROSE-CD: Robust One-step Speech Enhancement via Consistency Distillation, a novel approach for distilling a one-step consistency model. Specifically, we introduce a randomized learning trajectory to improve the model's robustness to noise. Furthermore, we jointly optimize the one-step model with two time-domain auxiliary losses, enabling it to recover from teacher-induced errors and surpass the teacher model in overall performance. This is the first pure one-step consistency distillation model for diffusion-based speech enhancement, achieving 54 times faster inference speed and superior performance compared to its 30-step teacher model. Experiments on the VoiceBank-DEMAND dataset demonstrate that the proposed model achieves state-of-the-art performance in terms of speech quality. Moreover, its generalization ability is validated on both an out-of-domain dataset and real-world noisy recordings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。