针对语音增强模型压缩难题,提出动态频率自适应知识蒸馏方法。
Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement
- 根据频段特性动态调整蒸馏目标,区分高低频学习需求。
- 在三种先进模型上压缩后性能优于传统蒸馏方法。
- 适合部署在计算资源受限设备上的语音增强模型优化。
基于深度学习的语音增强(SE)模型近期超越了传统方法,但其在资源受限设备上的部署仍面临高计算和内存开销的挑战。本文提出一种新型动态频率自适应知识蒸馏(DFKD)方法,有效压缩SE模型。该方法动态评估模型输出,区分高低频成分,并针对不同频段的特性自适应调整学习目标,充分利用SE任务的内在特征。为验证DFKD的有效性,我们在三种先进模型(DCCRN、ConTasNet、DPTNet)上进行了实验。结果表明,该方法不仅显著提升压缩后模型(学生模型)的性能,且在针对SE任务的基于logit的知识蒸馏方法中表现更优。
原文摘要 · Abstract (English)
Deep learning-based speech enhancement (SE) models have recently outperformed traditional techniques, yet their deployment on resource-constrained devices remains challenging due to high computational and memory demands. This paper introduces a novel dynamic frequency-adaptive knowledge distillation (DFKD) approach to effectively compress SE models. Our method dynamically assesses the model's output, distinguishing between high and low-frequency components, and adapts the learning objectives to meet the unique requirements of different frequency bands, capitalizing on the SE task's inherent characteristics. To evaluate the DFKD's efficacy, we conducted experiments on three state-of-the-art models: DCCRN, ConTasNet, and DPTNet. The results demonstrate that our method not only significantly enhances the performance of the compressed model (student model) but also surpasses other logit-based knowledge distillation methods specifically for SE tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。