通过双侧依赖控制,实现专家模型路由的协同优化与负载均衡。
Hierarchical Copula-Gumbel-Top-K Routing: Two-Sided Dependence Control for Frozen Mixture-of-Experts at Fixed Per-Token Routing Laws
- 引入分层耦合-古贝尔-顶K路由,控制不同令牌间的依赖关系。
- 正向组内耦合提升专家选择一致性,反向跨组耦合降低负载方差。
- 冻结主模型下仅训练小控制器,适合高效微调场景。
一个随机的古贝尔-顶K路由为混合专家(MoE)模型中每个令牌定义了一个路由规则:即在专家列表和混合权重上的分布。我们探讨在保持每个令牌完整路由规则不变的前提下,不同令牌间可达到的联合路由分布有哪些。提出一种双侧构造方法——分层耦合-古贝尔-顶K(CGA)。在相关令牌组内,使用交换对称的高斯耦合正向关联各专家坐标上的古贝尔扰动,增强组内专家集的一致性;在互不重叠的组对之间,采用可调节的反向构造引入可控的负相关性。证明两种操作均保持每个令牌的有序顶K采样、混合权重及包含概率分布与独立路由相同;因此条件期望专家流量也得以保留。我们刻画了此机制的权衡:组内正向耦合只能增加实际专家负载的方差,而跨组非负反向耦合只能减少该方差,相较于同强度组内耦合时的平坦耦合。一致性与负载分散由此由两个互补的依赖调节器共同控制。由于基础模型保持不变,调节器可由小控制器驱动,通过得分函数估计器训练:主模型仅需前向传播,梯度仅作用于控制器。初步小规模实验验证了机制与训练路径的有效性,但尚未建立任务级微调收益。
原文摘要 · Abstract (English)
A stochastic Gumbel-Top-K router defines, for every token of a mixture-of-experts (MoE) model, a routing law: a distribution over ordered expert lists and mixture weights. We ask which joint distributions over the routing choices of different tokens are reachable while every individual token's complete routing law is held exactly fixed. We give a two-sided construction, Hierarchical Copula-Gumbel-Top-K (CGA). Within a group of related tokens, an exchangeable Gaussian copula positively correlates the Gumbel perturbations at each expert coordinate, which can increase within-group expert-set coherence. Across disjoint pairs of groups, a tunable antithetic construction introduces a selectable amount of negative dependence. We prove that both operations leave each token's ordered Top-K sample, mixture weights, and inclusion probabilities identical in distribution to independent routing at a routing layer conditioned on its pre-routing logits; conditional expected expert traffic is preserved as a consequence. We characterize the resulting trade-off: positive within-group coupling can only inflate the variance of realized expert loads relative to independent routing, while nonnegative cross-group opposition can only reduce it relative to flat coupling at the same within-group strength. Coherence and load dispersion are thus controlled by two complementary dependence dials on the invariance constraint surface. Because the base model is untouched, the dials can be driven by a small controller over frozen features, trainable with a score-function estimator: the frozen network is evaluated only in the forward direction, and gradients are confined to the controller. An initial small-scale pilot validates the mechanism and the training route, but does not establish task-level fine-tuning gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。