提出新方法提升大模型强化学习中的推理能力
Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR

- 用条件互信息分析每个词元的奖励响应机制
- 发现高熵词元在持续推理中贡献更大
- 适合研究大模型强化学习与推理优化的人
强化学习结合可验证奖励(RLVR)能提升大语言模型的推理能力,但稀疏的结果奖励使得词元级别的信用分配困难。本文将词元级别的信用视为行为策略到事后后验的奖励条件转移。在自回归RLVR中,这一转移可通过条件互信息(CMI)表达,表明词元熵对可能的回顾性信用有上界约束。然而,熵仅反映容量而非更新方向,因此引入四象限分解,按奖励极性和词元熵分离更新。控制干预实验显示,这两者共同决定词元更新。持续推理收益集中在符号为正且高熵的象限,而低熵更新则快速饱和。基于此,提出前瞻性感知策略优化(HAPO),一种保持符号的GRPO改进方法,实现容量引导的优势重分配。在两种模型设置下的数学推理基准测试中,HAPO表现优于其他熵感知基线。
原文摘要 · Abstract (English)
Reinforcement Learning with Verifiable Rewards (RLVR) improves the reasoning ability of Large Language Models (LLMs), but sparse outcome rewards make token-level credit assignment difficult. We study token-level credit as a reward-conditioned shift from the behavior policy to a hindsight posterior. In autoregressive RLVR, this shift can be expressed through Conditional Mutual Information (CMI), which shows that token entropy upper-bounds possible hindsight credit. Entropy, however, indicates capacity rather than update direction, so we introduce the Four Quadrant Decomposition to separate updates by reward polarity and token entropy. Controlled interventions show that these two factors jointly shape token updates. Sustained reasoning gains concentrate in signed high-entropy quadrants, whereas low-entropy updates saturate quickly. Based on this analysis, we propose Hindsight-Aware Policy Optimization (HAPO), a sign-preserving modification to GRPO that performs capacity-guided advantage reallocation. Experiments on mathematical reasoning benchmarks in two model settings show that HAPO achieves competitive performance among entropy-aware baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。