通过有限专家池分析通信高效MoE的路由信息,揭示了模型泛化与路由效率的关系。
Expert Routing for Communication-Efficient MoE via Finite Expert Banks
- 构建基于预训练CNN的有限专家池,用数据依赖规则选择专家
- 发现估计的互信息 $\hat{I}(S;W)$ 随泛化误差单调变化
- 提供可计算的 $I(X;T)$ 和速率-精度曲线工具,适合系统设计者
资源高效的机器学习越来越多地采用稀疏混合专家(MoE)架构,其中门控机制既是学习组件,也作为控制计算、通信和准确率的路由接口。受有限速率下门控解释的启发,我们将门控视为随机信道,并使用 $I(X;T)$ 量化选定专家可获得的路由信息。为使相关信息量在真实数据上可处理,我们构建了一个基于预训练CNN专家和离散数据依赖选择规则的有限专家池MNIST。由于所选模型属于有限候选集,算法互信息 $I(S;W)$ 可通过经验后验 $q(W|S)$ 得到闭式离散熵估计。扫描数据依赖参数 $α$,我们观察到 $\hat{I}(S;W)$ 单调追踪泛化差距,而Xu-Raginsky界表现出预期的松散性。我们还与均匀联合界基线比较,并引入 $I(X;T)$ 的经验估计器,结合Blahut-Arimoto算法绘制专家库上的准确率-速率曲线。该框架为资源感知的MoE推理系统分析提供了实用工具,可将 $I(X;T)$ 与 $D(R_g)$ 作为高效专家路由的设计代理。
原文摘要 · Abstract (English)
Resource-efficient machine learning increasingly uses sparse Mixture-of-Experts (MoE) architectures, where the gate acts as both a learning component and a routing interface controlling computation, communication, and accuracy. Motivated by finite-rate interpretations of MoE gating, we treat the gate as a stochastic channel and use $I(X;T)$ to quantify the routing information available to the selected expert. To make the associated information quantities tractable beyond synthetic examples, we develop a finite-bank MNIST construction using pretrained CNN experts and a discrete, data-dependent selection rule. Since the selected model belongs to a finite candidate set, the algorithmic mutual information $I(S;W)$ admits a closed-form discrete-entropy estimator from the empirical posterior $q(W|S)$. Sweeping a data-dependence parameter $α$, we observe that $\widehat I(S;W)$ monotonically tracks the generalization gap, while the Xu-Raginsky bound exhibits the expected looseness. We also compare with a uniform union-bound baseline and introduce an empirical estimator of $I(X;T)$ together with a Blahut-Arimoto procedure for tracing an accuracy-rate curve over the expert bank. The proposed framework provides a practical tool for analyzing resource-aware MoE inference systems and for interpreting $I(X;T)$ and $D(R_g)$ as design proxies for efficient expert routing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。