arXiv:2505.03442cs.SDcs.LG2025-05

用余弦相似度对齐隐空间,让小模型更好复现大模型的降噪能力

Knowledge Distillation for Speech Denoising by Latent Representation Alignment with Cosine Distance

  • 通过余弦相似度对齐教师与学生模型的隐层表示
  • 在多种不匹配场景下,学生模型性能优于基线方法
  • 适合资源受限设备部署,如助听器、智能眼镜

语音降噪是广泛应用的重要任务,但现有高性能方法过于复杂,难以在手机、智能眼镜、助听器等低资源设备上部署。知识蒸馏可通过将复杂教师模型的知识迁移到轻量学生模型来缓解这一问题。然而,现有方法常因限制学生学习教师的分布、信息顺序和特征维度而影响效果。本文提出一种新方法,利用去噪自编码框架、线性逆瓶颈结构及余弦相似性特性,实现隐空间表示对齐。在公开数据集上进行多组跨模型不匹配实验,报告了所提方法与当前最优基线方法在指标上的均值与标准差。结果表明,该方法使学生模型性能更优,且能容忍更强的师生模型差异。

原文摘要 · Abstract (English)

Speech denoising is a generally adopted and impactful task, appearing in many common and everyday-life use cases. Although there are very powerful methods published, most of those are too complex for deployment in everyday and low-resources computational environments, like hand-held devices, intelligent glasses, hearing aids, etc. Knowledge distillation (KD) is a prominent way for alleviating this complexity mismatch and is based on the transferring/distilling of knowledge from a pre-trained complex model, the teacher, to another less complex one, the student. Existing KD methods for speech denoising are based on processes that potentially hamper the KD by bounding the learning of the student to the distribution, information ordering, and feature dimensionality learned by the teacher. In this paper, we present and assess a method that tries to treat this issue, by exploiting the well-known denoising-autoencoder framework, the linear inverted bottlenecks, and the properties of the cosine similarity. We use a public dataset and conduct repeated experiments with different mismatching scenarios between the teacher and the student, reporting the mean and standard deviation of the metrics of our method and another, state-of-the-art method that is used as a baseline. Our results show that with the proposed method, the student can perform better and can also retain greater mismatching conditions compared to the teacher.

语音降噪知识蒸馏隐空间对齐轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。