arXiv:2602.19345cs.LGcs.AI2026-02

用平滑门控函数提升大模型策略优化的稳定性与性能

Smooth Gate Functions for Soft Advantage Policy Optimization

  • 采用平滑门控函数替代硬截断,提升训练稳定性
  • 实验显示新门控函数在数学推理任务上性能更优
  • 为大模型策略优化提供可复用的稳定设计思路

组相对策略优化(GRPO)显著提升了大语言模型的训练效果和推理能力,但其依赖硬截断导致训练不稳定。软自适应策略优化(SAPO)通过引入基于sigmoid的平滑门控函数解决了该问题,实现更稳定的更新。本文进一步探索不同门控函数对训练稳定性和最终模型性能的影响,形式化了有效门控应具备的关键性质,并识别出多个可用于实证评估的函数族。基于Qwen2.5-7B-Instruct模型在数学推理任务上的实验结果,为设计更平滑、更鲁棒的大语言模型策略优化目标提供了实用指导。

原文摘要 · Abstract (English)

Group Relative Policy Optimization (GRPO) has significantly advanced the training of large language models and enhanced their reasoning capabilities, while it remains susceptible to instability due to the use of hard clipping. Soft Adaptive Policy Optimization (SAPO) addresses this limitation by replacing clipping with a smooth sigmoid-based gate function, which leads to more stable updates. We have decided to push this theory further and investigate the impact of different gate functions on both training stability and final model performance. We formalize the key properties that admissible gates should satisfy and identify several families of such functions for empirical evaluation. This paper presents an analysis of our findings based on experiments conducted with the Qwen2.5-7B-Instruct model on mathematical reasoning tasks. These results provide practical guidance for designing smoother and more robust policy optimization objectives for large language model training.

策略优化大模型训练平滑门控Qwen

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。