arXiv:2502.17187cs.CLcs.AI2025-02被引 1

分析MoE大模型在答题任务中专家的贡献,发现多数专家从未被激活。

Evaluating Expert Contributions in a MoE LLM for Quiz-Based Tasks

  • 通过答题基准测试评估每个专家的激活情况
  • 多数专家在推理中未被激活,门控网络分布接近均匀
  • 同一层内专家平均表现差异显著,提示能力不均

近期,采用专家混合(MoE)结构的大语言模型受到广泛关注,当前最先进的模型普遍采用此架构。尽管已有大量研究聚焦于模型训练和超参数选择,但对MoE层性质的后评估分析仍较为匮乏。本文首次尝试填补这一空白,基于问答类的MMLU基准对MoE层中专家的贡献进行评估。结果表明,在该基准测试中,大多数专家在推理过程中从未被激活;此外,门控网络的输出分布更接近均匀分布而非稀疏分布;最后,我们发现同一层中部分专家的平均表现存在显著差异。

原文摘要 · Abstract (English)

Recently, Large Language Models (LLMs) with Mixture of Experts (MoE) layers have gained significant attention. Currently, state-of-the-art LLMs utilize this architecture. There is a substantial amount of research on how to train such models and how to select hyperparameters for this architecture. However, there is a lack of studies focusing on post-evaluation analysis of MoE layer properties. In this paper, we take a first step toward closing this gap by evaluating expert contributions on the quiz-based MMLU benchmark. We show that most experts were never activated during inference on this benchmark. Additionally, the output distribution of gating networks is much closer to uniform than sparse. Finally, we demonstrate that the average performance of some experts within the same layer varies significantly.

MoE大模型评估专家贡献

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。