arXiv:2606.14398cs.LG2026-06

用离散语言模型解释专家分工,揭示任务特化机制。

A theoretical model for task routing in mixture-of-expert transformers

论文配图:A theoretical model for task routing in mixture-of-expert transformers
图 1 · 摘自论文原文
  • 构建基于语法模板和词典的离散语言模型
  • 证明单层MoE可按任务复杂度分配专用专家
  • 为模型中局部知识电路提供理论依据

混合专家(MoE)层可在保持推理计算量不变的前提下扩展Transformer模型。尽管前沿MoE模型中已观察到任务-专家专业化现象,但现有理论工作依赖连续混合模型,难以有效建模自然语言。一个关键未解问题是:如何用离散语言模型从理论上解释Transformer MoE中的任务专业化?为此,我们通过语法模板和有限键值词典表示结构化知识,并形式化证明单层MoE Transformer可通过专用于特定任务的专家编码知识。该构造表明,查询会被路由至唯一、任务特定的专家,其规模仅取决于任务内在复杂度(即对应语法模板与事实词典的总大小)。该构造为MoE模型中观察到的局部知识电路提供了理论支持。我们通过实验验证了不同MoE损失函数下的模型性能,进一步支持理论发现。

原文摘要 · Abstract (English)

Mixture-of-experts (MoE) layers enable the scaling of transformer models while keeping the inference compute fixed. While task-expert specialization has been observed in empirical studies of frontier MoE transformer models, existing theoretical work analyzes this using continuous mixture models that cannot be used to model natural language effectively. An important open question is to \textit{theoretically explain task-expert specialization in transformer MoE models using discrete models of language}. To address this, we represent structured knowledge via syntactic templates and finite key-value dictionaries, and prove formally that a single-layer MoE transformer can encode knowledge by using experts that specialize in the corresponding tasks. Our construction shows how queries are routed to unique, task-specific experts whose size depends solely on the intrinsic complexity of the given task (i.e. the combined size of its syntactic templates and factual dictionary). Our construction provides a theoretical support for empirical results on localized knowledge circuits in MoE models. We support our theoretical findings with experiments evaluating model performance under varying MoE loss functions.

MoE专家分工理论分析语言建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。