通过分语言教师蒸馏,提升多语言语音识别模型的性能与专长。
Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR

- 分语言教师独立优化,再通过路由和逐标记蒸馏融合知识。
- 在中英粤及方言上超越所有教师,平均提升3.2%词错误率。
- 适合追求高精度多语言语音识别的开发者与研究者。
现代基于大语言模型的多语言语音识别系统已将多语言能力作为标准功能,利用大规模多语言语料和大语言模型的跨语言知识,在多语言基准测试中表现优异。然而,由于不同语言在声学、音系和词汇特征上的异质性,联合建模会引入优化冲突,削弱各语言的专长。为此,我们提出语言专用多教师在线蒸馏(LS-MOPD),将语言专有知识获取与多语言能力整合解耦:语言专用教师通过强化学习独立优化,其专长再通过语言路由和标记级多教师蒸馏融入通用多语言学生模型,从而减少直接跨语言优化冲突。我们进一步探索了静态与动态声学前缀配置,考察教师-学生前缀一致性对在线蒸馏效果的影响。在涵盖普通话、方言、粤语和英语的基准测试上,实验表明LS-MOPD显著优于强化学习基线,并在几乎所有基准上超越最佳教师的实证性能上限,揭示其在多语言语音识别中具备超越所有教师的泛化潜力。
原文摘要 · Abstract (English)
Modern LLM-based ASR systems have established multilingual capability as a standard feature, leveraging large-scale multilingual corpora and LLMs' cross-lingual knowledge to achieve competitive performance across multilingual benchmarks. However, jointly modeling languages with heterogeneous acoustic, phonological, and lexical characteristics inevitably introduces optimization conflicts, undermining language-wise specialization. To address this challenge, we propose Language-Specialized Multi-Teacher On-Policy Distillation (LS-MOPD), which decouples language-specific knowledge acquisition from multilingual capability integration: language-specialized teachers are independently optimized via reinforcement learning (RL), with their expertise then integrated into a generalist multilingual student through language routing and token-level multi-teacher distillation, thereby reducing direct cross-lingual optimization conflicts. We further explore static and dynamic acoustic-prefix configurations to examine how teacher-student prefix consistency influences the efficacy of on-policy distillation. Experiments on benchmarks covering Mandarin, Mandarin subdialects, Cantonese, and English demonstrate that LS-MOPD substantially outperforms RL baselines and surpasses the empirical performance envelope defined by the best-performing RL teachers on nearly all benchmarks, revealing its potential to generalize beyond all teachers in multilingual ASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。