提升多语言语音识别在跨领域场景下的鲁棒性
BLR-MoE: Boosted Language-Routing Mixture of Experts for Domain-Robust Multilingual E2E ASR
- 将专家路由机制扩展至自注意力层,缓解语言混淆
- 通过专家剪枝与路由器增强,提升路由准确性
- 在1万小时数据集上验证,显著改善跨领域表现
近年来,混合专家(MoE)架构如语言路由MoE(LR-MoE)被广泛用于缓解多语言自动语音识别(MASR)中的语言混淆问题。然而,在跨领域场景下仍存在显著的语言混淆现象。本文将LR-MoE中的语言混淆解耦为自注意力模块和路由模块的混淆。为缓解自注意力中的混淆,提出在MASR中引入注意力-混合专家(attention-MoE)结构,使MoE同时作用于前馈网络(FFN)和自注意力层。此外,为增强基于语言识别(LID)的路由机制对语言混淆的鲁棒性,提出专家剪枝与路由增强方法。结合上述改进,构建了增强型语言路由混合专家(BLR-MoE)架构。在10,000小时的多语言语音识别数据集上验证了其有效性。
原文摘要 · Abstract (English)
Recently, the Mixture of Expert (MoE) architecture, such as LR-MoE, is often used to alleviate the impact of language confusion on the multilingual ASR (MASR) task. However, it still faces language confusion issues, especially in mismatched domain scenarios. In this paper, we decouple language confusion in LR-MoE into confusion in self-attention and router. To alleviate the language confusion in self-attention, based on LR-MoE, we propose to apply attention-MoE architecture for MASR. In our new architecture, MoE is utilized not only on feed-forward network (FFN) but also on self-attention. In addition, to improve the robustness of the LID-based router on language confusion, we propose expert pruning and router augmentation methods. Combining the above, we get the boosted language-routing MoE (BLR-MoE) architecture. We verify the effectiveness of the proposed BLR-MoE in a 10,000-hour MASR dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。