arXiv:2607.24030cs.CLcs.SD2026-07中稿 · COLM

通过分组语言专家,让多语种语音识别更高效且准确。

MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition

论文配图:MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition
图 1 · 摘自论文原文
  • 将相似语言分组,用专属专家模块提升识别效率。
  • 在495种语言上表现优于传统模型,参数增加极少。
  • 适合需要大规模多语种语音识别的场景。

覆盖数百种语言的大规模多语种自动语音识别(ASR)模型需在多样语言和声学条件下保持稳健性能。然而,多语种困境常导致模型容量被稀释。为此,我们提出语言分组专家混合模型(MoLGE),基于语音自监督模型(S3M)。MoLGE为相似语言集群分配专用专家模块,相比传统语言特定的专家混合(MoE)方案减少了子模块数量。同时,在S3M架构的解耦声学与语言组件中引入分层低秩适配(LoRA)策略,实现语言特异性建模的同时保持参数高效。进一步研究了基于语言学与数据驱动的标准对语言分组的影响,提供了可解释的视角,揭示语言结构如何影响多语种系统的可扩展性。在涵盖495种语言的多语种基准上评估,结果表明MoLGE始终优于密集型多语种基线,仅小幅增加可训练参数。值得注意的是,这些语言分组策略显著提升了发音与拼写层面的建模效果。研究发现,结构化语言专属性是实现大规模多语种语音识别的有效路径。

原文摘要 · Abstract (English)

Massively multilingual automatic speech recognition (ASR) models covering hundreds of languages must maintain robust performance across diverse linguistic and acoustic conditions. However, these models often encounter the curse of multilinguality, where model capacity is diluted across languages. To address this challenge, we propose Mixture of Language Group Experts (MoLGE), built upon speech self-supervised models (S3Ms). MoLGE assigns dedicated expert modules to clusters of similar languages, reducing the number of required submodules compared to conventional language-specific Mixture-of-Experts (MoE) schemes. It further integrates a hierarchical Low-Rank Adaptation (LoRA) strategy into the disentangled acoustic and linguistic components of the S3M architecture, enabling efficient modeling of language-specific characteristics while maintaining parameter efficiency. Further, we investigate the impact of language grouping strategies based on both linguistic and data-driven criteria on overall performance, providing an interpretable perspective on how language structure influences scalability in multilingual speech systems. In experiments, we evaluate MoLGE on a multilingual benchmark encompassing 495 languages. Results demonstrate that MoLGE consistently outperforms dense multilingual baselines with a minimal increase in trainable parameters. Notably, these language grouping strategies yield substantial improvements for both phonetic and orthographic aspects of ASR modeling. Our findings suggest that structured language specialization provides an effective pathway for massively scaling language coverage of multilingual ASR.

多语种识别专家混合语音识别参数效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。