arXiv:2604.03592cs.CLcs.AI2026-04中稿 · EMNLP被引 4

发现多语言MoE模型中语言路由隔离现象,提出针对性优化框架

Unveiling Language Routing Isolation in Multilingual MoE Models for Interpretable Subnetwork Adaptation

  • 通过分析专家路由模式,发现高低资源语言激活不同专家集
  • 提出RISE框架,仅训练特定专家子网,低资源语言F1提升最高10.85%
  • 适合需要提升特定语言性能且不希望影响其他语言的研究者

混合专家(MoE)模型在不同语言间表现差异显著,但其内部机制尚不明确。本文系统分析了MoE模型中的专家路由模式,揭示了一种称为语言路由隔离的现象:高资源与低资源语言倾向于激活几乎不重叠的专家集合。分层分析显示,路由模式随模型深度呈现逐层收敛-发散特征。基于此,我们提出RISE(路由隔离引导的子网络增强)框架,利用路由隔离识别并优化语言特定专家子网。RISE采用三重选择策略:用特异性分数筛选浅层和深层的语言专属专家,用重叠分数选择中间层的通用专家。仅训练选定子网而冻结其余参数,显著提升低资源语言性能,同时保持其他语言能力。在10种语言上的实验表明,RISE可实现目标语言最高10.85%的F1提升,跨语言退化极小。

原文摘要 · Abstract (English)

Mixture-of-Experts (MoE) models exhibit striking performance disparities across languages, yet the internal mechanisms driving these gaps remain poorly understood. In this work, we conduct a systematic analysis of expert routing patterns in MoE models, revealing a phenomenon we term Language Routing Isolation, in which high- and low-resource languages tend to activate largely disjoint expert sets. Through layer-stratified analysis, we further show that routing patterns exhibit a layer-wise convergence-divergence pattern across model depth. Building on these findings, we propose RISE (Routing Isolation-guided Subnetwork Enhancement), a framework that exploits routing isolation to identify and adapt language-specific expert subnetworks. RISE applies a tripartite selection strategy, using specificity scores to identify language-specific experts in shallow and deep layers and overlap scores to select universal experts in middle layers. By training only the selected subnetwork while freezing all other parameters, RISE substantially improves low-resource language performance while preserving capabilities in other languages. Experiments on 10 languages demonstrate that RISE achieves target-language F1 gains of up to 10.85% with minimal cross-lingual degradation.

多语言MoE专家路由子网络优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。