arXiv:2506.01503cs.LGcs.SD2025-06中稿 · Interspeech 2025

提出对齐空白符的新方法,让语音模型压缩无需标签数据

Analyzing the Importance of Blank for CTC-Based Knowledge Distillation

  • 设计对称选择机制,动态优化空白符处理方式
  • 移除CTC损失后识别准确率仅下降0.8%,接近原性能
  • 可实现无标签音频上的知识蒸馏,适用于海量未标注语音

随着大规模预训练语音模型的兴起,其推理延迟和成本显著增加。为保留模型性能的同时提升效率,知识蒸馏成为重要手段。本文研究基于CTC的蒸馏变体,重点关注空白符(blank token)的处理策略。发现常见的空白符消除方法在实际应用中并不总有效。为此,我们探索新的空白符选择模式,找到标准知识蒸馏与空白符消除之间的平衡点。通过引入对称选择方法,成功在知识蒸馏阶段完全移除CTC损失,性能下降仅0.8%(在LibriSpeech test-clean上),几乎无损。该方法使训练过程摆脱目标标签依赖,有望在无转录音频数据上实现知识蒸馏。

原文摘要 · Abstract (English)

With the rise of large pre-trained foundation models for automatic speech recognition new challenges appear. While the performance of these models is good, runtime and cost of inference increases. One approach to make use of their strength while retaining efficiency is to distill their knowledge to smaller models during training. In this work, we explore different CTC-based distillation variants, focusing on blank token handling. We show that common approaches like blank elimination do not always work off the shelf. We explore new blank selection patterns as a potential sweet spot between standard knowledge distillation and blank elimination mechanisms. Through the introduction of a symmetric selection method, we are able to remove the CTC loss during knowledge distillation with minimal to no performance degradation. With this, we make the training independent from target labels, potentially allowing for distillation on untranscribed audio data.

语音识别知识蒸馏CTC无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。