arXiv:2602.15091stat.MLcs.IT2026-02被引 2

研究专家模型在通信受限下的表现极限,揭示了通信速率与泛化能力的权衡关系。

Mixture-of-Experts under Finite-Rate Gating: Communication--Generalization Trade-offs

  • 将门控机制建模为有限信息速率的信道,用信息论分析其性能边界
  • 推导出门控速率、表达能力和泛化误差之间的定量权衡关系
  • 适用于需要高效通信的大型专家模型系统设计,如边缘计算场景

Mixture-of-Experts(MoE)架构通过门控机制选择专用专家子网络来分解预测任务。本文从信息论视角分析门控机制,将门控建模为在有限信息速率下运行的随机信道。基于信息论学习框架,我们推导了互信息形式的泛化界,并构建了有限速率门控的率失真特性 $D(R_g)$,其中 $R_g := I(X; T)$。在标准经验率失真最优条件下,得到期望风险上界:${mathbb{E}[R(W)]} \le D(R_g)+δ_m+\sqrt{(2/m){ I(S; W)}}$。该分析揭示了通信受限下MoE系统的容量感知限制。在合成多专家模型上的数值模拟验证了门控速率、表达能力与泛化性能之间的预期权衡。

原文摘要 · Abstract (English)

Mixture-of-Experts (MoE) architectures decompose prediction tasks into specialized expert sub-networks selected by a gating mechanism. This letter adopts a communication-theoretic view of MoE gating, modeling the gate as a stochastic channel operating under a finite information rate. Within an information-theoretic learning framework, {we specialize a mutual-information generalization bound and develop a rate-distortion characterization $D(R_g)$ of finite-rate gating, where $R_g:=I(X; T)$, yielding (under a standard empirical rate-distortion optimality condition) $\mathbb{E}[R(W)] \le D(R_g)+δ_m+\sqrt{(2/m)\, I(S; W)}$. }The analysis yields capacity-aware limits for communication-constrained MoE systems, and numerical simulations on synthetic multi-expert models empirically confirm the predicted trade-offs between gating rate, expressivity, and generalization.

MoE信息论通信效率泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。