arXiv:2605.15403cs.LGmath.OC2026-05被引 4

提出ϕ-平衡框架,让专家模型更稳定高效地分配任务。

$ϕ$-Balancing for Mixture-of-Experts Training

论文配图:$ϕ$-Balancing for Mixture-of-Experts Training
图 1 · 摘自论文原文
  • 基于期望路由分布的凸函数优化,直接追求全局平衡
  • 在大规模预训练和微调中均超越现有方法,专家利用率更均衡
  • 算法轻量,仅需极小开销,适合工业级MoE系统

Mixture-of-Experts(MoE)模型依赖于专家使用的均衡性以充分发挥其可扩展性。然而,现有的负载平衡方法多为启发式,且基于有噪声的小批量分配统计,与总体目标存在偏差。本文提出ϕ-平衡,一种基于严格凸、对称且可微的势函数的原理性框架,直接最小化预期路由分布的势能,从而实现群体层面的专家平衡。利用凸对偶性,我们推导出等价的极小极大形式,并通过镜面下降获得一个简洁的在线算法,实现低开销的EMA路由调整。在大规模预训练和下游微调中,ϕ-平衡始终优于先前的Switch风格和无损失基线,展现出更稳定、更有效的专家使用。

原文摘要 · Abstract (English)

Mixture-of-Experts (MoE) models rely on balanced expert utilization to fully realize their scalability. However, existing load-balancing methods are largely heuristic and operate on noisy mini-batch assignment statistics, introducing bias relative to population-level objectives. We propose $ϕ$-balancing, a principled framework that directly targets population-level expert balance by minimizing a strictly convex, symmetric, and differentiable potential of the expected routing distribution. Using convex duality, we derive an equivalent min-max formulation and obtain a simple online algorithm via mirror descent, yielding an efficient EMA-based routing adjustment with negligible overhead. Across large-scale pretraining and downstream fine-tuning, $ϕ$-balancing consistently outperforms prior Switch-style and loss-free baselines, demonstrating more stable and effective expert utilization.

MoE专家模型负载均衡凸优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。