arXiv:2605.28139cs.AI2026-05

用少量语音数据让小模型通过教师指导提升语音识别能力

Data-Efficient On-Policy Distillation for Automatic Speech Recognition

论文配图:Data-Efficient On-Policy Distillation for Automatic Speech Recognition
图 1 · 摘自论文原文
  • 用教师模型指导学生模型在真实语音上进行在线蒸馏
  • 仅用10万小时语音即超越同规模基线,接近17亿参数模型表现
  • 适合资源有限但想提升小模型性能的研究者和开发者

构建高性能语音识别模型通常依赖大规模音频标注数据,导致复现与定制成本高昂。本文研究了基于10万小时语音训练的0.6B参数音频条件语言模型Ark-ASR,探索强教师模型Qwen-ASR能否通过在线蒸馏传递额外识别能力。在中英文语音识别基准上,所提训练方法持续优于纯监督微调,且在五个评估集中的四个超越同规模的Qwen3-ASR-0.6B基线。该效果仅需10万小时语音,远低于Qwen3-Omni AuT编码器所需的2000万小时。尽管更大的Qwen3-ASR-1.7B仍更优,结果表明教师引导的在线训练可在极低语音预算下显著缩小紧凑模型与大模型的差距。支持重叠诊断进一步表明,教师数据阶段提升了学生与教师的局部匹配度,与近期关于在线蒸馏有效性的分析一致。

原文摘要 · Abstract (English)

Building competitive automatic speech recognition (ASR) models usually requires large-scale au- dio supervision, which makes reproduction and specialization expensive. We study Ark-ASR, a 0.6B- parameter audio-conditioned language model trained with 100k hours of speech, and examine whether a strong Qwen-ASR teacher can transfer additional recognition capability through on-policy distillation. Across Mandarin and English ASR benchmarks, the proposed training recipe consistently improves over supervised fine-tuning alone and outperforms the same-scale Qwen3-ASR-0.6B baseline on four of five evaluation sets. This is achieved with only 100k hours of speech, compared with the 20M hours of super- vised audio reported for the Qwen3-Omni AuT encoder. The larger Qwen3-ASR-1.7B remains stronger, but the results show that teacher-guided on-policy training can substantially close the gap for compact ASR models under a much smaller audio budget. A support-overlap diagnostic further suggests that the teacher-data stage improves local student-teacher compatibility, matching recent analyses of when on-policy distillation is effective.

语音识别知识蒸馏小样本学习模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。