arXiv:2601.16117cs.SDcs.CV2026-01中稿 · ICASSP 2026

通过知识蒸馏提升动态语音模型的性能与效率

Distillation-based Layer Dropping (DLD): Effective End-to-end Framework for Dynamic Speech Networks

  • 用知识蒸馏优化层跳过策略,实现端到端动态调整
  • 高跳过率下词错误率降低9.32%,训练时间减少33.3%
  • 适合资源受限设备上的实时语音识别应用

边缘设备在资源受限且多变的环境下运行,需要能够自适应计算资源的动态模型。为满足这一需求,层跳过($/mathcal{LD}$)方法常被用来将静态模型转变为动态模型,通过跳过网络部分结构来降低计算复杂度。然而,现有$/mathcal{LD}$方法在低和高跳过率情况下均显著影响模型性能,破坏了性能与计算量之间的平衡。为此,我们提出一种基于知识蒸馏的层跳过(DLD)框架,以端到端方式有效结合知识蒸馏与$/mathcal{LD}$能力,显著提升动态语音网络的表现。在三个公开语音识别基准上,使用Conformer和WavLM等知名模型进行的全面实验表明,该框架在高跳过和无跳过场景下分别将词错误率降低9.32%和2.25%,同时训练时间减少33.3%。

原文摘要 · Abstract (English)

Edge devices operate in constrained and varying resource settings, requiring dynamic architectures that can adapt to limitations of the available resources. To meet such demands, layer dropping ($\mathcal{LD}$) approach is typically used to transform static models into dynamic ones by skipping parts of the network along with reducing overall computational complexity. However, existing $\mathcal{LD}$ methods greatly impact the dynamic model's performance for low and high dropping cases, deteriorating the performance-computation trade-off. To this end, we propose a distillation-based layer dropping (DLD) framework that effectively combines the capabilities of knowledge distillation and $\mathcal{LD}$ in an end-to-end fashion, thereby achieving state-of-the-art performance for dynamic speech networks. Comprehensive experimentation utilizing well-known speech recognition methods, including conformer and WavLM, on three public benchmarks demonstrates the effectiveness of our framework, reducing the word error rate by $9.32\%$ and $2.25\%$ for high and no dropping cases with $33.3\%$ reduction in training time.

语音识别动态网络知识蒸馏边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。