用动态专家路由提升多语言语音识别效果,避免专家失效。
Dynamic Multi-Expert Projectors with Stabilized Routing for Multilingual Speech Recognition
- 设计动态多专家投影模块,稳定分配不同语言的声学特征映射。
- 在4种印度语言上实现最高7.6%的词错误率降低,性能显著提升。
- 专家路由展现语言相关性,适合多语言语音系统研发者参考。
基于大语言模型的语音识别将冻结的语音编码器与大型语言模型通过轻量级投影模块连接。尽管在单语场景下有效,单一投影模块难以捕捉多语言语音识别所需的多样化声学到语义映射。为此,我们提出SMEAR-MoE,一种具备稳定路由机制的混合专家投影模块,确保所有专家都能获得密集梯度,防止专家崩溃,同时支持跨语言共享。我们在四种印地语系语言(印地语、马拉地语、泰米尔语、泰卢固语)上系统比较了单体、静态多投影和动态混合专家结构。SMEAR-MoE表现优异,相比单投影基线最多实现7.6%的相对词错误率(WER)下降,且运行效率相当。专家路由分析显示,语义相关的语言共享相同专家,说明其具有语言层面的专属性。结果表明,稳定的多专家投影是实现可扩展、鲁棒多语言语音识别的关键。
原文摘要 · Abstract (English)
Recent advances in LLM-based ASR connect frozen speech encoders with Large Language Models (LLMs) via lightweight projectors. While effective in monolingual settings, a single projector struggles to capture the diverse acoustic-to-semantic mappings required for multilingual ASR. To address this, we propose SMEAR-MoE, a stabilized Mixture-of-Experts projector that ensures dense gradient flow to all experts, preventing expert collapse while enabling cross-lingual sharing. We systematically compare monolithic, static multi-projector, and dynamic MoE designs across four Indic languages (Hindi, Marathi, Tamil, Telugu). Our SMEAR-MoE achieves strong performance, delivering upto a 7.6% relative WER reduction over the single-projector baseline, while maintaining comparable runtime efficiency. Analysis of expert routing further shows linguistically meaningful specialization, with related languages sharing experts. These results demonstrate that stable multi-expert projectors are key to scalable and robust multilingual ASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。