揭示MoE语言模型中领域专家与驱动专家的激活规律
What Gets Activated: Uncovering Domain and Driver Experts in MoE Language Models
- 区分领域与驱动专家,用熵和因果指标定位激活模式
- 早期词元更易触发驱动专家,影响模型输出决策
- 调节两类专家权重可显著提升三类模型性能
多数可解释性研究聚焦于Transformer的层或神经元级机制,而对MoE大模型中的专家层级行为关注不足。受人脑功能特化的启发,本文分析了MoE模型在三个公开领域的专家激活情况,回答两个核心问题:(1)哪些专家被激活,特定类型专家是否呈现稳定激活模式;(2)词元如何关联并触发特定专家激活。为此,提出基于熵和因果效应的度量方法,分别识别领域专家与驱动专家。进一步探究词元与专家激活的关系。结果表明:(1)激活专家中部分具有明显领域偏好,另一些则对模型表现有强因果影响,起决定性作用;(2)句子中较早出现的词元更易触发驱动专家;(3)调整领域与驱动专家权重,在所有三类模型与领域中均带来显著性能提升。研究揭示了MoE模型内部机制,提升了其可解释性。
原文摘要 · Abstract (English)
Most interpretability work focuses on layer- or neuron-level mechanisms in Transformers, leaving expert-level behavior in MoE LLMs underexplored. Motivated by functional specialization in the human brain, we analyze expert activation by distinguishing domain and driver experts. In this work, we study expert activation in MoE models across three public domains and address two key questions: (1) which experts are activated, and whether certain expert types exhibit consistent activation patterns; and (2) how tokens are associated with and trigger the activation of specific experts. To answer these questions, we introduce entropy-based and causal-effect metrics to assess whether an expert is strongly favored for a particular domain, and how strongly expert activation contributes causally to the model's output, thus identify domain and driver experts, respectively. Furthermore, we explore how individual tokens are associated with the activation of specific experts. Our analysis reveals that (1) Among the activated experts, some show clear domain preferences, while others exert strong causal influence on model performance, underscoring their decisive roles. (2) tokens occurring earlier in a sentence are more likely to trigger the driver experts, and (3) adjusting the weights of domain and driver experts leads to significant performance gains across all three models and domains. These findings shed light on the internal mechanisms of MoE models and enhance their interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。