arXiv:2605.04066cs.CLcs.ET2026-05ACL

动态调整奖励策略,让大模型推理更准更快

Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning

论文配图:Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning
图 1 · 摘自论文原文
  • 用可变均值机制自动调节学习强度
  • 数学推理任务平均得分提升3.0分
  • 适合追求高精度推理的AI研究者

强化学习结合可验证奖励(RLVR)是提升大语言模型(LLM)推理能力的关键范式。然而,现有方法多采用静态策略优化,难以匹配模型推理能力的动态变化。为此,我们提出自适应幂均策略优化(APMPO),包含两大创新:幂均策略优化(PMPO)和反馈自适应裁剪(FAC)。PMPO引入广义幂均目标,使模型能自适应地在算术平均的信号放大与几何平均的一致性增强之间切换;FAC基于实时奖励统计动态调整裁剪范围,克服静态机制局限。依托这些改进,APMPO显著优化了学习动态与推理性能。在三个推理任务、九个数据集上的实验表明,其优于当前最优的基于RLVR的方法。例如,在使用Qwen2.5-3B-Instruct时,数学推理基准的Pass@1平均得分比GRPO高出3.0分。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Rewards (RLVR) is an essential paradigm that enhances the reasoning capabilities of Large Language Models (LLMs). However, existing methods typically rely on static policy optimization schemes that misalign with the model's evolving reasoning capabilities. To address this issue, we propose Adaptive Power-Mean Policy Optimization (APMPO), which comprises two main innovations: Power-Mean Policy Optimization (PMPO) and Feedback-Adaptive Clipping (FAC). Specifically, PMPO introduces a generalized power-mean objective. This enables the model to adaptively transition from the signal-amplifying behavior of the arithmetic mean to the consistency-enforcing behavior of the geometric mean. FAC adaptively adjusts clipping bounds based on real-time reward statistics to overcome the limitations of static mechanisms. Capitalizing on these innovations, APMPO improves learning dynamics and reasoning performance. Extensive experiments on nine datasets across three reasoning tasks showcase the superiority of APMPO over state-of-the-art RLVR-based baselines. For instance, APMPO boosts the average Pass@1 score on mathematical reasoning benchmarks by 3.0 points compared to GRPO when using Qwen2.5-3B-Instruct.

强化学习大模型推理自适应优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。