arXiv:2409.15905cs.SDcs.AI2024-09被引 13

用专家混合模型提升多语种语音识别效果

Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM

  • 设计插入删除中断标记机制,增强大语言模型的文本生成能力
  • 采用专家混合架构连接器,高效管理多语言语音特征映射
  • 两阶段渐进训练策略,兼顾专业与通用表征学习

本文提出一种融合专家混合(MoE)结构的语音条件大语言模型(LLM),以应对自动语音识别中的多语种切换(Code-Switching, CS)挑战。我们设计了插入与删除中断标记(IDIT)机制,有效提升大语言模型在语音识别任务中的文本生成能力。同时,提出基于MoE的连接器架构,可高效管理多语言特征表示。进一步地,采用两阶段渐进训练策略:第一阶段解冻连接器并使用语言专用专家,将语音表征映射至文本空间;第二阶段激活所有专家,联合训练连接器与LLM LoRA适配器,学习通用表征。实验表明,该方法显著优于现有先进模型,包括端到端及大规模音视频语言模型。

原文摘要 · Abstract (English)

In this paper, we introduce a speech-conditioned Large Language Model (LLM) integrated with a Mixture of Experts (MoE) based connector to address the challenge of Code-Switching (CS) in Automatic Speech Recognition (ASR). Specifically, we propose an Insertion and Deletion of Interruption Token (IDIT) mechanism for better transfer text generation ability of LLM to speech recognition task. We also present a connecter with MoE architecture that manages multiple languages efficiently. To further enhance the collaboration of multiple experts and leverage the understanding capabilities of LLM, we propose a two-stage progressive training strategy: 1) The connector is unfrozen and trained with language-specialized experts to map speech representations to the text space. 2) The connector and LLM LoRA adaptor are trained with the proposed IDIT mechanism and all experts are activated to learn general representations. Experimental results demonstrate that our method significantly outperforms state-of-the-art models, including end-to-end and large-scale audio-language models.

语音识别多语种专家混合大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。