arXiv:2601.00557cs.CLcs.SD2026-01被引 1

轻量多语言语音识别模型,无需语言标签也能高效识别。

A Language-Agnostic Hierarchical LoRA-MoE Architecture for CTC-based Multilingual ASR

  • 分层LoRA-MoE架构,共享与专用专家协同学习
  • 在两个数据集上推理延迟降低8.2%~11.7%
  • 适合资源受限设备部署的无标签多语言识别

大规模多语言语音识别(mASR)模型如Whisper虽性能优异,但计算和延迟开销大,限制了在资源受限边缘设备上的部署。本文提出一种基于CTC架构的轻量级、语言无关的多语言语音识别系统,结合领域自适应技术。具体地,将语言无关分层LoRA-MoE(HLoRA)框架集成至mHuBERT-CTC模型中,通过语言识别后验驱动的LoRA路由实现端到端解码。其分层设计包含多语言共享LoRA以学习语言不变声学表征,以及语言特异的LoRA专家以建模语言依赖特性。所提路由机制在推理时无需事先语言身份信息或显式语言标签,实现真正的语言无关解码。在MSR-86K和MLC-SLM 2025 Challenge数据集上的实验表明,HLoRA性能接近两阶段推理方法,同时分别降低实时因子(RTF)11.7%和8.2%,显著提升低资源mASR应用的解码效率。

原文摘要 · Abstract (English)

Large-scale multilingual ASR (mASR) models such as Whisper achieve strong performance but incur high computational and latency costs, limiting their deployment on resource-constrained edge devices. In this study, we propose a lightweight and language-agnostic multilingual ASR system based on a CTC architecture with domain adaptation. Specifically, we introduce a Language-agnostic Hierarchical LoRA-MoE (HLoRA) framework integrated into an mHuBERT-CTC model, enabling end-to-end decoding via LID-posterior-driven LoRA routing. The hierarchical design consists of a multilingual shared LoRA for learning language-invariant acoustic representations and language-specific LoRA experts for modeling language-dependent characteristics. The proposed routing mechanism removes the need for prior language identity information or explicit language labels during inference, achieving true language-agnostic decoding. Experiments on MSR-86K and the MLC-SLM 2025 Challenge datasets demonstrate that HLoRA achieves comparable performance to two-stage inference approaches while reducing RTF by 11.7% and 8.2%, respectively, leading to improved decoding efficiency for low-resource mASR applications.

语音识别多语言LoRA边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。