arXiv:2510.23027cs.LGcs.CL2025-10ACL被引 5

提出新方法提升MoE模型强化学习训练稳定性与效果

Towards Stable and Effective Reinforcement Learning for Mixture-of-Experts

  • 基于路由器输出设计重缩放策略,降低梯度方差
  • 使MoE模型在强化学习中收敛更稳定,性能显著提升
  • 适合研究大规模专家模型高效训练的学者参考

近年来,强化学习(RL)的进展大幅提升了大规模语言模型的训练效果,显著增强了生成质量和推理能力。然而,现有研究大多聚焦于密集模型,而对混合专家(MoE)架构的RL训练仍缺乏深入探索。为解决MoE训练中常见的不稳定性问题,本文提出一种新型路由器感知的方法,用于优化离策略强化学习中的重要性采样权重。具体而言,我们设计了一种由路由器logits引导的重缩放策略,有效降低了梯度方差,缓解了训练发散问题。实验结果表明,该方法显著提升了MoE模型的收敛稳定性与最终性能,凸显了针对MoE架构量身定制强化学习算法的潜力,并为大规模专家模型的高效训练提供了新方向。

原文摘要 · Abstract (English)

Recent advances in reinforcement learning (RL) have substantially improved the training of large-scale language models, leading to significant gains in generation quality and reasoning ability. However, most existing research focuses on dense models, while RL training for Mixture-of-Experts (MoE) architectures remains underexplored. To address the instability commonly observed in MoE training, we propose a novel router-aware approach to optimize importance sampling (IS) weights in off-policy RL. Specifically, we design a rescaling strategy guided by router logits, which effectively reduces gradient variance and mitigates training divergence. Experimental results demonstrate that our method significantly improves both the convergence stability and the final performance of MoE models, highlighting the potential of RL algorithmic innovations tailored to MoE architectures and providing a promising direction for efficient training of large-scale expert models.

强化学习MoE训练稳定大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。