arXiv:2606.10439cs.SDcs.CL2026-06中稿 · ICASSP 2026被引 1

用专家混合与动态下采样提升多语言语音识别精度

Enhancing Multilingual LLM-based ASR with Mixture of Experts and Dynamic Downsampling

论文配图:Enhancing Multilingual LLM-based ASR with Mixture of Experts and Dynamic Downsampling
图 1 · 摘自论文原文
  • 引入专家混合架构增强跨语言适应能力
  • 通过连续积分-放电机制实现动态下采样与模态对齐
  • 适合关注多语言语音模型优化的研究者

大规模语言模型(LLM)的快速发展为自动语音识别(ASR)开辟了新方向,其有效集成成为关键且具有挑战性的研究课题。本文提出一种基于投影器的LLM-ASR框架,针对多语言泛化和模态对齐两大核心挑战。方法上结合专家混合(MoE)架构以提升跨语言适应性,并采用连续积分-放电(CIF)机制实现动态下采样与模态对齐。实验结果表明,该组合显著提升性能,超越多个强基线模型。所提方法为构建更准确、鲁棒且通用的基于LLM的ASR系统迈出重要一步。

原文摘要 · Abstract (English)

The rapid progress of large language models (LLMs) has opened up a new frontier for automatic speech recognition (ASR), making their effective integration a critical and challenging research direction. To this end, this work proposes a projector-based LLM-ASR framework targeting the key challenges of multilingual generalization and modality alignment. Our approach incorporates a Mixture of Experts (MoE) architecture to improve cross-lingual adaptability, and a Continuous Integrate-and-Fire (CIF) mechanism for dynamic downsampling and modality alignment. Experimental results show that the combination of these components yields substantial performance improvements, surpassing strong baseline models. The proposed method represents a step toward building more accurate, robust, and generalizable LLM-based ASR systems.

语音识别多语言专家混合模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。