arXiv:2505.16991cs.CV2025-05中稿 · InterSpeech 2025被引 2

从大模型高效生成小模型,训练更快且更准。

An Effective Training Framework for Light-Weight Automatic Speech Recognition Models

  • 两阶段表征学习,用大模型指导小模型训练
  • 训练速度提升三倍,词错误率降低最多12.54%
  • 适合资源受限设备部署的轻量语音识别模型

深度学习的进步推动了大型自动语音识别(ASR)模型的发展,尽管性能优异,但其计算和内存开销限制了在低资源设备上的部署。现有方法(如剪枝、知识蒸馏、跳过层等)虽能将大模型压缩为小模型,却常导致性能大幅下降,或需长时间训练小模型才能获得较好效果。为此,我们提出一种高效的两阶段表征学习框架,仅需一次训练即可从一个大模型生成多个小型模型,在有限训练轮次内实现显著更优的性能。在多个ASR基准测试上进行的全面实验表明,该方法实现了三倍的训练加速,并在词错误率(WER)上最高提升12.54%。

原文摘要 · Abstract (English)

Recent advancement in deep learning encouraged developing large automatic speech recognition (ASR) models that achieve promising results while ignoring computational and memory constraints. However, deploying such models on low resource devices is impractical despite of their favorable performance. Existing approaches (pruning, distillation, layer skip etc.) transform the large models into smaller ones at the cost of significant performance degradation or require prolonged training of smaller models for better performance. To address these issues, we introduce an efficacious two-step representation learning based approach capable of producing several small sized models from a single large model ensuring considerably better performance in limited number of epochs. Comprehensive experimentation on ASR benchmarks reveals the efficacy of our approach, achieving three-fold training speed-up and up to 12.54% word error rate improvement.

语音识别轻量模型训练加速表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。