arXiv:2409.02050cs.CLcs.SD2024-09中稿 · IEEE SLT 2024被引 3

用语言识别引导专家协作,提升多语混杂语音识别准确率

Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model

  • 通过语言识别权重引导专家选择与协作
  • 在多个数据集上显著优于现有方法
  • 保持MoE模型高效推理特性,无需额外预训练

由于跨语言语音特征的固有相似性,多语混杂语音识别面临巨大挑战。本文提出协同式混合专家(Collaborative-MoE)模型,利用专家组间的协同机制。首先,前置路由网络显式学习语言识别(LID)任务,并基于获得的LID权重选择专家,确保MoE层具备鲁棒的路由信息,减少不同语言域对专家参数更新的干扰。同时,利用LID权重促进组间协作,实现语言特异性表征融合。此外,在每个语言专家组内,门控网络无监督地促进语言之外属性的协作。大量实验表明,该方法显著优于其他方案,且保持MoE模型高效的推理能力,无需额外预训练。

原文摘要 · Abstract (English)

Due to the inherent difficulty in modeling phonetic similarities across different languages, code-switching speech recognition presents a formidable challenge. This study proposes a Collaborative-MoE, a Mixture of Experts (MoE) model that leverages a collaborative mechanism among expert groups. Initially, a preceding routing network explicitly learns Language Identification (LID) tasks and selects experts based on acquired LID weights. This process ensures robust routing information to the MoE layer, mitigating interference from diverse language domains on expert network parameter updates. The LID weights are also employed to facilitate inter-group collaboration, enabling the integration of language-specific representations. Furthermore, within each language expert group, a gating network operates unsupervised to foster collaboration on attributes beyond language. Extensive experiments demonstrate the efficacy of our approach, achieving significant performance enhancements compared to alternative methods. Importantly, our method preserves the efficient inference capabilities characteristic of MoE models without necessitating additional pre-training.

语音识别多语混杂MoE模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。