arXiv:2603.16929cs.LGcs.AI2026-03

MHPO通过动态调节策略更新,提升强化学习训练稳定性。

MHPO: Modulated Hazard-aware Policy Optimization for Stable Reinforcement Learning

  • 引入可微分的对数保真调制器,约束重要性比率范围。
  • 设计解耦风险惩罚项,分别控制策略正负方向偏移。
  • 适用于需要稳定训练的文本与视觉语言任务,防崩溃、防退化。

在基于组相对策略优化(GRPO)的框架中,调控重要性比率对训练稳定性至关重要。然而,现有方法如硬截断存在不可微边界和梯度消失问题,难以保持梯度准确性,且缺乏对极端偏差的自适应抑制机制,使优化过程易受突发策略跳变影响。为此,我们提出调制型风险感知策略优化(MHPO),一种鲁棒稳定的强化学习新框架。MHPO引入对数保真调制器(LFM),将无界的重要性比率映射至有界可微域,有效防止高方差异常词元破坏损失景观,并保障全局梯度稳定。同时,解耦风险惩罚(DHP)结合生存分析中的累积风险函数,独立调节正负向策略更新。通过风险感知惩罚塑造优化景观,实现对不对称策略偏移的细粒度调控,同时缓解过度扩张导致的模式坍塌和灾难性收缩引发的策略退化,在稳定信任域内运行。在涵盖文本与视觉语言任务的多种推理基准上,大量实验表明,MHPO始终优于现有方法,性能更优且训练稳定性显著提升。

原文摘要 · Abstract (English)

Regulating the importance ratio is critical for the training stability of Group Relative Policy Optimization (GRPO) based frameworks. However, prevailing ratio control methods, such as hard clipping, suffer from non-differentiable boundaries and vanishing gradient regions, failing to maintain gradient fidelity. Furthermore, these methods lack a hazard-aware mechanism to adaptively suppress extreme deviations, leaving the optimization process vulnerable to abrupt policy shifts. To address these challenges, we propose Modulated Hazard-aware Policy Optimization (MHPO), a novel framework designed for robust and stable reinforcement learning. The proposed MHPO introduces a Log-Fidelity Modulator (LFM) to map unbounded importance ratios into a bounded, differentiable domain. This mechanism effectively prevents high-variance outlier tokens from destabilizing the loss landscape while ensuring global gradient stability. Complementarily, a Decoupled Hazard Penalty (DHP) integrates cumulative hazard functions from survival analysis to independently regulate positive and negative policy shifts. By shaping the optimization landscape with hazard-aware penalties, the proposed MHPO achieves fine-grained regulation of asymmetric policy shifts simultaneously mitigating mode collapse from over-expansion and preventing policy erosion from catastrophic contraction within a stabilized trust region. Extensive evaluations on diverse reasoning benchmarks across both text-based and vision-language tasks demonstrate that MHPO consistently outperforms existing methods, achieving superior performance while significantly enhancing training stability.

强化学习策略优化训练稳定风险感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。