用知识蒸馏指导结构化剪枝,让语音说话人分离模型更小更快无损
Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models
- 通过知识蒸馏引导的结构化剪枝压缩自监督模型
- 模型体积减少80%,推理速度提升4倍,性能不降
- 在多个数据集上泛化能力强,无需领域适配
自监督学习模型如WavLM通过提供丰富的上下文语音表征,显著提升了说话人分离性能。然而,这些模型的高计算和内存开销限制了其在实时及资源受限场景中的部署。本文系统研究了基于知识蒸馏的结构化剪枝在压缩基于自监督学习的分离模型中的应用。我们探讨了针对模型参数与计算复杂度的剪枝目标,并分析了不同策略,发现一种简单的整体剪枝方法在效率与精度之间取得了最佳平衡。所提方法实现了最高达80%的模型尺寸缩减和4倍的推理加速,且性能无下降。在八个公开的分离数据集上的综合实验表明,剪枝模型始终匹配或优于未剪枝版本。此外,在CHiME-6数据集上展现出强域外泛化能力,无需任何领域适应即可达到CHiME-7挑战赛顶尖系统水平的准确率。结果表明,经蒸馏引导的结构化剪枝可生成高效且泛化能力强的分离系统,适用于真实应用场景。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) models such as WavLM have substantially advanced speaker diarization by providing rich contextual speech representations. However, the high computational and memory costs of these models hinder deployment in real-time and resource-constrained scenarios. This work presents a systematic study on compressing SSL-based diarization models through structured pruning guided by knowledge distillation. We investigate pruning objectives that target both model parameters and computational complexity, and analyze alternative strategies, showing that a simple overall pruning approach provides the best balance between efficiency and accuracy. Our method achieves up to 80% model size reduction and 4x faster inference without performance degradation. Comprehensive experiments across eight public diarization datasets demonstrate that the pruned models consistently match or surpass the performance of their uncompressed counterparts. Furthermore, we show strong out-of-domain generalization on the CHiME-6 dataset, achieving accuracy comparable to the top systems in the CHiME-7 challenge without any domain adaptation. These results highlight that structured pruning, when guided by distillation, can yield efficient and generalizable diarization systems suitable for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。