arXiv:2602.01103cs.AI2026-02中稿 · ICML被引 2

揭示MoE模型强化学习不稳定的根源,提出目标级劫持新视角

Probing RLVR training instability through the lens of objective-level hacking

  • 从目标级劫持出发,分析奖励信号在词元层面的信用错配机制
  • 发现训练-推理差异异常增长是核心病理性动态,且有明确因果解释
  • 适用于研究大模型强化学习稳定性的算法设计者与架构开发者

长时间基于可验证奖励的强化学习(RLVR)已被证明能持续提升大语言模型的推理能力,但其训练常出现不稳定性,尤其在混合专家(MoE)架构中更为显著。这种不稳定性严重阻碍模型能力提升,但其根本原因和机制仍不清楚。本文提出一种系统性框架,通过目标级劫持的视角理解RLVR的不稳定性。不同于由可被利用的验证器引发的奖励劫持,目标级劫持源于词元级别的信用分配错位,并表现为优化目标中的系统级虚假信号。基于该框架,结合对300亿参数MoE模型的广泛实验,我们追溯并形式化了MoE模型中一种关键病理性训练动态的起源:训练-推理差异的异常增长。这一现象此前虽与不稳定性密切相关,但缺乏机制解释。本研究提供了对MoE模型训练不稳定性背后动态过程的明确、因果性说明,为设计稳定可靠的RLVR算法提供指导。

原文摘要 · Abstract (English)

Prolonged reinforcement learning with verifiable rewards (RLVR) has been shown to drive continuous improvements in the reasoning capabilities of large language models, but the training is often prone to instabilities, especially in Mixture-of-Experts (MoE) architectures. Training instability severely undermines model capability improvement, yet its underlying causes and mechanisms remain poorly understood. In this work, we introduce a principled framework for understanding RLVR instability through the lens of objective-level hacking. Unlike reward hacking, which arises from exploitable verifiers, objective-level hacking emerges from token-level credit misalignment and is manifested as system-level spurious signals in the optimization objective. Grounded in our framework, together with extensive experiments on a 30B MoE model, we trace the origin and formalize the mechanism behind a key pathological training dynamic in MoE models: the abnormal growth of the training-inference discrepancy, a phenomenon widely associated with instability but previously lacking a mechanistic explanation. These findings provide a concrete and causal account of the training dynamics underlying instabilities in MoE models, offering guidance for the design of stable RLVR algorithms.

强化学习MoE架构训练稳定大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。