arXiv:2502.05837eess.AS2025-02被引 1

将知识蒸馏与结构化剪枝结合,显著提升自监督语音模型性能。

Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models

  • 融合知识蒸馏与稀疏正则化剪枝,协同优化语音模型。
  • 非流式场景下相对错误率降低8.9%,流式场景降13.4%。
  • 适合追求高精度语音识别的模型压缩研究者。

传统知识蒸馏用于模型压缩时常导致性能下降。本文评估了将知识蒸馏损失与低秩分解(LRF)及l0正则化等剪枝技术结合,在基于Conformer的自监督学习预训练网络上的效果。还提出一种联合剪枝与训练RNN-T语音识别模型的方法,结果表明该方法优于先剪枝再训练的流程。该策略在非流式场景中实现8.9%相对词错误率(RWER)改进,流式场景达13.4%改进,均优于基线。

原文摘要 · Abstract (English)

Traditionally, Knowledge Distillation (KD) is used for model compression, often leading to suboptimal performance. In this paper, we evaluate the impact of combining KD loss with alternative pruning techniques, including Low-Rank Factorization (LRF) and l0 regularization, on a conformer-based pre-trained network under the paradigm of Self-Supervised Learning (SSL). We also propose a strategy to jointly prune and train an RNN-T-based ASR model, demonstrating that this approach yields superior performance compared to pruning a pre-trained network first and then using it for ASR training. This approach led to a significant reduction in word error rate: l0 and KD combination achieves the best non-streaming performance, with a 8.9% Relative Word Error Rate (RWER) improvement over the baseline, while LRF and KD combination yields the best results for streaming ASR, improving RWER by 13.4%.

语音识别模型压缩知识蒸馏剪枝

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。