共享路由决策提升语音识别中稀疏专家模型的协作与性能
Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition
- 跨层共享路由器,增强不同层专家间的协同
- 在10个外部数据集上平均降低11.2%的词错误率
- 适合需要高效、鲁棒语音识别的工业级应用
混合专家(MoE)架构已从语言建模扩展至自动语音识别(ASR)。传统MoE方法如Switch Transformer在每层内独立路由专家。我们分析发现,多数层的路由器选择的专家与其它层关联性弱。为增强不同层专家间的协作并促进专业化,我们在不同MoE层间使用共享路由器,提出Omni-router Transformer。在大规模伪标签数据集上的大量实验以及在10个多样化的域外ASR基准测试中的评估表明,Omni-router Transformer能实现更低训练损失,持续优于密集模型和Switch Transformer,在平均词错误率上分别降低11.2%和8.2%,同时提供结构化专家使用方式并增强对多样化数据的鲁棒性。
原文摘要 · Abstract (English)
Mixture-of-experts (MoE) architectures have expanded from language modeling to automatic speech recognition (ASR). Traditional MoE methods, such as the Switch Transformer, route experts independently within each layer. Our analysis reveals that routers in most layers make expert choices that are not strongly correlated with the choices of the routers in other layers. To increase the cooperation between experts in different layers and encourage greater specialization, we use a shared router across different MoE layers. We call this model Omni-router Transformer. Extensive experiments on a large-scale pseudo-labeled dataset and evaluations across 10 diverse, out-of-domain ASR benchmarks demonstrate that the Omni-router Transformer is able to achieve lower training loss and consistently outperform dense and Switch Transformer models, reducing average word error rates by 11.2% and 8.2%, respectively, while providing structured expert usage and improved robustness to diverse data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。