arXiv:2506.13127cs.SDeess.AS2025-06

通过时频校准蒸馏,提升语音增强中轻量模型的性能。

Leveraging Local and Global Knowledge Integration with Time-Frequency Calibrated Distillation for Speech Enhancement

  • 构建跨集与集内递归融合框架,利用时频差异信息进行知识蒸馏。
  • 多层交互蒸馏在时频域分别计算相似性权重,实现精准分配蒸馏贡献。
  • 适用于低复杂度学生模型,在单/多通道数据集上均优于现有方案。

本文提出一种基于时频校准蒸馏的集内与集间递归融合框架(I²SRF-TFCKD)用于语音增强(SE)。不同于以往蒸馏方法,该框架充分挖掘语音的时频差异特性,同时兼顾局部信息聚焦与全局知识传播。首先,构建协同蒸馏范式以捕捉集内与集间相关性:在相关集合内,多层师生特征成对匹配并进行校准蒸馏;随后通过递归融合生成代表性特征,形成融合特征集以支持集间知识交互。其次,提出基于双流时频交叉校准的多层交互蒸馏策略,分别在时域与频域计算师生相似性校准权重,并进行交叉加权,从而根据语音特性精细化分配各层蒸馏贡献。该蒸馏策略应用于在 L3DAS23 挑战赛语音增强赛道排名第一的双路径空洞卷积循环网络(DPDCRN)。在单通道与多通道语音增强数据集上的客观评估表明,所提蒸馏策略能持续有效提升轻量级学生模型性能,优于其他蒸馏方案。

原文摘要 · Abstract (English)

In this paper, we propose an intra-set and inter-set recursive fusion framework with time-frequency calibrated knowledge distillation (I$^2$SRF-TFCKD) for SE. Different from previous distillation strategies for SE, the proposed framework fully exploits the time-frequency differential information of speech while facilitating both local information focusing and global knowledge circulation. Firstly, we construct a collaborative distillation paradigm for intra-set and inter-set correlations. Within a correlated set, multi-layer teacher-student features are pairwise matched for calibrated distillation. Subsequently, we generate representative features from each correlated set through recursive fusion to form the fused feature set that enables inter-set knowledge interaction. Secondly, we propose a multi-layer interactive distillation based on dual-stream time-frequency cross-calibration, which calculates the teacher-student similarity calibration weights in the time and frequency domains respectively and performs cross-weighting, thus enabling refined allocation of distillation contributions across different layers according to speech characteristics. The proposed distillation strategy is applied to the dual-path dilated convolutional recurrent network (DPDCRN) that ranked first in the SE track of the L3DAS23 challenge. To evaluate the effectiveness of I$^2$SRF-TFCKD, we conduct experiments on both single-channel and multi-channel SE datasets. Objective evaluations demonstrate that the proposed KD strategy consistently and effectively improves the performance of the low-complexity student model and outperforms other distillation schemes.

语音增强知识蒸馏时频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。