揭示MoE模型中稀疏性如何提升可解释性
Sparsity and Superposition in Mixture of Experts
- 用网络稀疏度替代特征稀疏度,重新定义专家分工机制
- 稀疏性越高,专家越专注单一语义,实现更强的单义性
- 适合关注模型可解释性与性能平衡的研究者
混合专家(MoE)模型已成为大规模语言模型扩展的核心,但其机制与密集网络的差异仍不明确。以往研究指出,密集模型通过超叠加(superposition)在少于维度的条件下表示更多特征,且该现象依赖于特征稀疏性和重要性。然而,MoE模型无法用相同框架解释。我们发现,特征稀疏性和重要性不会引发突变相变,而网络稀疏性(活跃专家占比)更能刻画MoE特性。我们提出新指标衡量专家间的超叠加。结果表明,网络稀疏度越高,模型表现出更强的单义性(monosemanticity)。我们提出基于单义特征表示的专家专业化新定义,而非负载均衡。当初始化得当时,专家会自然聚集于连贯的特征组合。这表明,高网络稀疏性的MoE可在不牺牲性能的前提下实现更高可解释性,挑战了可解释性与能力不可兼得的普遍假设。
原文摘要 · Abstract (English)
Mixture of Experts (MoE) models have become central to scaling large language models, yet their mechanistic differences from dense networks remain poorly understood. Previous work has explored how dense models use \textit{superposition} to represent more features than dimensions, and how superposition is a function of feature sparsity and feature importance. MoE models cannot be explained mechanistically through the same lens. We find that neither feature sparsity nor feature importance cause discontinuous phase changes, and that network sparsity (the ratio of active to total experts) better characterizes MoEs. We develop new metrics for measuring superposition across experts. Our findings demonstrate that models with greater network sparsity exhibit greater \emph{monosemanticity}. We propose a new definition of expert specialization based on monosemantic feature representation rather than load balancing, showing that experts naturally organize around coherent feature combinations when initialized appropriately. These results suggest that network sparsity in MoEs may enable more interpretable models without sacrificing performance, challenging the common assumption that interpretability and capability are fundamentally at odds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。