arXiv:2409.11214eess.AScs.SD2024-09被引 13

双编码器+语言适配连接器,提升多语言语音转写效果

Ideal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text

  • 用双语编码器融合Whisper与MMS特征,增强多语言表征
  • 语言适配连接器使各语言误差率降低32.6%,翻译得分达36.78
  • 适合多语言语音识别与翻译场景,尤其关注语言差异的改进

通过连接器将音频编码器与大语言模型(LLM)结合,使模型能够处理和理解音频模态,显著提升了语音转写任务(如自动语音识别ASR和自动语音翻译AST)的表现。然而,现有方法在多语言环境下常忽略语言适应性问题,仅依赖多语言数据而未充分应对语言差异。为此,我们提出Ideal-LLM模型,采用双语多语言编码器以丰富语言特征信息,并引入语言适配连接器,针对每种语言进行专门适配。该模型融合Whisper与MMS编码器的优势,实现更丰富的多语言表示。同时,语言适配连接器通过为每种语言定制的语言权重选择器,提升模态转换能力。实验表明,Ideal-LLM显著提升ASR性能,相比标准语音编码器与LLM集成方案,平均词错误率降低32.6%;在AST任务上达到平均BLEU分数36.78。

原文摘要 · Abstract (English)

Integrating audio encoders with LLMs through connectors has enabled these models to process and comprehend audio modalities, significantly enhancing speech-to-text tasks, including automatic speech recognition (ASR) and automatic speech translation (AST). However, these methods often overlook the critical aspect of language adaptation in multilingual settings, relying instead on multilingual data without adequately addressing language differences. To address this gap, we propose the Ideal-LLM model, which employs dual multilingual encoders to enrich language feature information and utilizes a language-adapted connector to target the adaptation of each language specifically. By leveraging the complementary strengths of Whisper and MMS encoders, our approach ensures richer multilingual representations. Additionally, the language-adapted connector enhances modal transformation via a language weight selector tailored for each language. Experimental results demonstrate that Ideal-LLM significantly improves ASR performance, achieving a 32.6% relative reduction in average word error rates compared to the standard speech encoder integrated with LLMs and yields an average BLEU score of 36.78 for AST task.

语音转写多语言大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。