arXiv:2602.19327cs.LGcs.AI2026-02

提出SSPO方法,提升大模型对齐训练的稳定性和效果。

Soft Sequence Policy Optimization

  • 用软门控机制融合词级别概率比,构建序列级重要性权重。
  • 在数学推理和编程任务中显著提升训练稳定性与性能。
  • 适合关注大模型对齐优化、强化学习训练效率的研究者。

近期大语言模型对齐研究多聚焦于基于组相对策略优化(GRPO)的新策略优化方法。主要方向包括:(i) 采用更契合任务中序列级奖励的序列级重要性采样权重;(ii) 探索替代PPO裁剪的方法,以避免训练信号损失和熵崩溃。本文提出软序列策略优化(SSPO),一种离策略强化学习目标,通过在序列级重要性权重中引入词级别概率比的软门控函数实现。我们提供了理论动机,并研究了改进优化行为的实际调整策略。实验表明,SSPO在数学推理和编程任务中均提升了训练稳定性和性能。

原文摘要 · Abstract (English)

A significant portion of recent research on Large Language Model (LLM) alignment focuses on developing new policy optimization methods based on Group Relative Policy Optimization (GRPO). Two prominent directions have emerged: (i) a shift toward sequence-level importance sampling weights that better align with the sequence-level rewards used in many tasks, and (ii) alternatives to the PPO-style clipping that aim to avoid the associated loss of training signal and entropy collapse. We introduce Soft Sequence Policy Optimization, an off-policy reinforcement learning objective that incorporates soft gating functions over token-level probability ratios within sequence-level importance weights. We provide theoretical motivation for SSPO and investigate practical modifications to improve optimization behavior. Empirically, we demonstrate that SSPO improves training stability and performance both in mathematical reasoning and coding tasks.

大模型对齐强化学习策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。