通过内部机制分析,揭示MoE模型的专家协作与神经元利用率规律。
Beyond Benchmarks: Understanding Mixture-of-Experts Models through Internal Mechanisms
- 引入MUI指标,量化专家激活与神经元利用率
- 发现模型演进中神经元利用率下降,反映更强泛化能力
- 适合研究MoE机制或优化模型效率的开发者
Mixture-of-Experts(MoE)架构因其推理时仅激活部分参数而具备高效与可扩展性,成为热门方向。然而现有研究多聚焦性能指标,对内部机制理解不足。本文通过引入内部度量MUI,显式分析路由机制与专家行为。系统评估多个公开MoE模型后发现:(1)随着模型演进,神经元利用率下降,反映更强泛化能力;(2)训练过程呈动态变化,仅依赖基准性能信号有限,而MUI揭示深层演化轨迹;(3)任务完成依赖多专家协同,共享专家推动注意力集中;(4)神经元级激活模式可作为数据多样性的细粒度代理。这些结果表明MUI是基准性能的有益补充,为理解MoE模型容量、动态特性与专业化提供新视角。项目主页:https://yingjiahao14.github.io/MoE-MUI/
原文摘要 · Abstract (English)
Mixture-of-Experts (MoE) architectures have emerged as a promising direction, offering efficiency and scalability by activating only a subset of parameters during inference. However, current research remains largely performance-centric, with limited understanding of its internal mechanisms, thereby constraining broader progress. In this work, we use an internal metric to investigate the mechanisms of MoE architecture by explicitly incorporating routing mechanisms and analyzing expert-level behaviors. Through systematic analyses of a wide range of publicly available MoE models, we uncover several findings: (1) neuron utilization decreases as models evolve, reflecting stronger generalization; (2) training exhibits a dynamic trajectory, where benchmark performance alone provides limited signal while MUI reveals deeper insights; (3) task completion emerges from collaborative contributions of multiple experts, with shared experts driving concentration; and (4) activation patterns at the neuron level provide a fine-grained proxy for data diversity. Together, these results demonstrate the potential of MUI as a complementary indicator to benchmark performance, offering new insights into the capacity, dynamics, and specialization of MoE models. Our project can be found at https://yingjiahao14.github.io/MoE-MUI/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。