发现大模型专家路由受输入语义影响,验证了语义驱动的专家选择机制。
Probing Semantic Routing in Large Mixture-of-Expert Models
- 通过控制输入语义变化,测试专家激活模式
- 相同目标词不同语义时专家重叠度显著下降
- 适合研究模型内部工作机制与语义理解的学者
过去一年,参数量超过1000亿的大规模混合专家(MoE)模型在开放领域日益普及。尽管其优势常被归因于效率,但已有研究探索了通过路由行为实现的功能分化。本文探究大型MoE模型中的专家路由是否受输入语义影响。为此,设计两项受控实验:首先,比较包含同一目标词但该词在不同语义下的句子对的专家激活;其次,在固定上下文条件下,用语义相近或相异的目标词进行替换。通过对比这些条件下的专家重叠情况,发现存在清晰且统计显著的语义路由证据。
原文摘要 · Abstract (English)
In the past year, large (>100B parameter) mixture-of-expert (MoE) models have become increasingly common in the open domain. While their advantages are often framed in terms of efficiency, prior work has also explored functional differentiation through routing behavior. We investigate whether expert routing in large MoE models is influenced by the semantics of the inputs. To test this, we design two controlled experiments. First, we compare activations on sentence pairs with a shared target word used in the same or different senses. Second, we fix context and substitute the target word with semantically similar or dissimilar alternatives. Comparing expert overlap across these conditions reveals clear, statistically significant evidence of semantic routing in large MoE models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。