arXiv:2603.11321cs.LGcs.AI2026-03被引 1

让失败变反馈,让模型从错误中自主学习。

Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings

  • 用事后锚定机制,只在失败时引入教师示范指导
  • 理论证明能渐进消除偏差,实现无偏梯度优化
  • 适合需要自适应训练的稀疏奖励场景

基于可验证奖励的强化学习(RLVR)已成为后训练推理模型的有前景范式。然而,群组方法如群组相对策略优化(GRPO)在稀疏奖励设置下面临关键困境:纯强化学习易出现优势坍塌和高方差梯度估计,而混合策略优化则引入持续分布偏差。为解决此矛盾,我们提出事后锚定策略优化(HAPO)。HAPO采用合成成功注入(SSI)算子,一种仅在失败时选择性锚定教师示范的回溯机制。该注入由受汤普森采样启发的门控机制控制,形成自主、自调节的课程。理论上,我们证明了HAPO具备渐近一致性:随着策略提升,教师信号自然衰减,从而恢复无偏的在线策略梯度。这确保了离线指导仅作为临时支架,而非持久上限,使模型能够突破静态教师强制的局限。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for post-training reasoning models. However, group-based methods such as Group Relative Policy Optimization (GRPO) face a critical dilemma in sparse-reward settings: pure Reinforcement Learning (RL) suffers from advantage collapse and high-variance gradient estimation, while mixed-policy optimization introduces persistent distributional bias. To resolve this dilemma, we introduce Hindsight-Anchored Policy Optimization (HAPO). HAPO employs the Synthetic Success Injection (SSI) operator, a hindsight mechanism that selectively anchors optimization to teacher demonstrations during failure. This injection is governed by a Thompson sampling-inspired gating mechanism, creating an autonomous, self-paced curriculum. Theoretically, we demonstrate that HAPO achieves \textit{asymptotic consistency}: by naturally annealing the teacher signal as the policy improves, HAPO recovers the unbiased on-policy gradient. This ensures off-policy guidance acts as a temporary scaffold rather than a persistent ceiling, enabling the model to surpass the limitations of static teacher forcing.

强化学习稀疏奖励自适应训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。