arXiv:2602.06453cs.LG2026-02被引 3

提出新方法提升大模型微调时的稳定性和推理能力。

On the Plasticity and Stability for Post-Training Large Language Models

  • 用贝叶斯框架将梯度视为随机变量,动态软投影解决冲突。
  • 训练轨迹更平滑,在多个推理任务上性能显著提升。
  • 适合需要稳定微调的大模型应用,如复杂推理系统。

训练稳定性仍是分组相对策略优化(GRPO)的关键瓶颈,常表现为推理灵活性与通用能力保持之间的权衡。我们识别出根本原因在于灵活性与稳定性梯度间的几何冲突,导致破坏性干扰。关键的是,我们指出确定性投影方法在GRPO中次优,因其忽略了基于组的梯度估计的内在随机性。为此,我们提出概率冲突解决(PCR),一个将梯度建模为随机变量的贝叶斯框架。PCR通过不确定性感知的‘软投影’机制动态仲裁冲突,优化信噪比。大量实验表明,PCR显著平滑了训练轨迹,并在多种推理任务中实现更优性能。

原文摘要 · Abstract (English)

Training stability remains a critical bottleneck for Group Relative Policy Optimization (GRPO), often manifesting as a trade-off between reasoning plasticity and general capability retention. We identify a root cause as the geometric conflict between plasticity and stability gradients, which leads to destructive interference. Crucially, we argue that deterministic projection methods are suboptimal for GRPO as they overlook the intrinsic stochasticity of group-based gradient estimates. To address this, we propose Probabilistic Conflict Resolution (PCR), a Bayesian framework that models gradients as random variables. PCR dynamically arbitrates conflicts via an uncertainty-aware ``soft projection'' mechanism, optimizing the signal-to-noise ratio. Extensive experiments demonstrate that PCR significantly smooths the training trajectory and achieves superior performance in various reasoning tasks.

大模型微调强化学习稳定性贝叶斯方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。