MoE模型在双语学习中自发形成语言结构化路由,无需顺序训练。
A Declarative-Procedural Perspective on Expert Routing in Bilingual Mixture-of-Experts Language Models

- 基于声明-程序框架,分析双语MoE模型的词法与句法路由模式。
- 无课程训练组在第5层互信息达0.2599,高于顺序训练组的0.1148。
- 分阶段双语训练可降低单一语言主导,实现更均衡的路由分布。
我们研究了混合专家(MoE)语言模型在双语语言习得过程中是否会发展出语言结构化的专家路由机制。受声明-程序框架启发,我们分析了一个在顺序语言暴露下训练的仅解码器英语-德语MoE Transformer中的词汇、语法和句法处理。我们构建了一个基于探测的验证集,并提取了逐标记的路由分布,使用互信息、路由熵和Jensen-Shannon距离量化不同语言类别下的专业化程度。课程训练模型在第5层达到0.1148的峰值互信息,表明不同语言类别间路由分布存在差异。令人意外的是,未采用课程设计的基线模型在混合英德语数据上训练,其聚合专业化程度更强,在同一层达到0.2599的峰值互信息。结果表明,即使没有顺序语言暴露,MoE路由模式中仍会出现可解释的语言组织。在第二个训练种子上的复现显示,无课程条件下的专业化集中于单一语言,且语言身份依赖训练种子;而课程训练始终产生稳定、语言平衡的路由特征;分阶段双语暴露并非简单提升专业化,而是降低了单一语言主导性。
原文摘要 · Abstract (English)
We investigate whether Mixture-of-Experts (MoE) language models develop linguistically structured expert routing during bilingual language acquisition. Inspired by the Declarative-Procedural framework, we analyze lexical, grammatical, and syntactic processing in a decoder-only English-German MoE Transformer trained under sequential language exposure. We construct a probe-based validation set and extract token-level routing distributions to quantify category-dependent specialisation using mutual information, routing entropy, and Jensen-Shannon distance. The curriculum-trained model exhibits a peak mutual information of 0.1148 at layer 5, indicating category-dependent differences in routing distributions across linguistic categories. Surprisingly, a no-curriculum baseline trained on mixed English-German data shows stronger aggregate specialisation, reaching a peak mutual information of 0.2599 at the same layer. These results suggest that interpretable linguistic organization emerges within MoE routing patterns even without sequential language exposure. A replication at a second training seed shows that the no-curriculum condition's specialisation concentrates on a single language whose identity is seed-dependent, whereas the curriculum consistently yields a stable, language-balanced routing profile; rather than uniformly increasing specialisation, staged bilingual exposure reduces single-language dominance. The official Github repository: https://github.com/Amrit828/DP-Theory-MOE-Interpretability-Research
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。