揭示MoE模型在复杂任务中的表达能力,突破维度与稀疏性限制。
On the Expressive Power of Mixture-of-Experts for Structured Complex Tasks
- 基于低维流形与稀疏结构,证明浅层MoE可克服维数灾难。
- 深层MoE以L层E专家实现指数级分段函数逼近,达E^L种结构化任务。
- 解析门控、专家数、层数等关键设计,指导MoE变体优化。
混合专家网络(MoEs)在现代深度学习中展现出卓越效率。尽管其经验表现优异,但其建模复杂任务的理论基础仍不清晰。本文系统研究了MoE在两类常见结构先验——低维性与稀疏性下的表达能力。对于浅层MoE,我们证明其能高效逼近定义在低维流形上的函数,克服维数灾难。对于深层MoE,我们表明具有L层和每层E个专家的MoE可逼近由E^L个分段组成的复合稀疏函数,实现指数级结构化任务数量。分析揭示了门控机制、专家网络、专家数量及层数等核心组件的作用,为MoE变体设计提供了自然建议。
原文摘要 · Abstract (English)
Mixture-of-experts networks (MoEs) have demonstrated remarkable efficiency in modern deep learning. Despite their empirical success, the theoretical foundations underlying their ability to model complex tasks remain poorly understood. In this work, we conduct a systematic study of the expressive power of MoEs in modeling complex tasks with two common structural priors: low-dimensionality and sparsity. For shallow MoEs, we prove that they can efficiently approximate functions supported on low-dimensional manifolds, overcoming the curse of dimensionality. For deep MoEs, we show that $\mathcal{O}(L)$-layer MoEs with $E$ experts per layer can approximate piecewise functions comprising $E^L$ pieces with compositional sparsity, i.e., they can exhibit an exponential number of structured tasks. Our analysis reveals the roles of critical architectural components and hyperparameters in MoEs, including the gating mechanism, expert networks, the number of experts, and the number of layers, and offers natural suggestions for MoE variants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。