arXiv:2602.05717cs.AI2026-02被引 8

提出新方法解决强化学习中探索崩溃问题,提升准确率与多样性。

Anchored Policy Optimization: Mitigating Exploration Collapse Via Support-Constrained Rectification

  • 用安全流形约束策略支持集,避免全局形状匹配
  • 在数学基准上实现更高通过率与更优解多样性
  • 适合追求高效且稳定探索的RL研究者

基于可验证奖励的强化学习(RLVR)被视为一种剪枝机制。我们识别出一种系统性病理——递归空间收缩(RSC),由正向锐化与负向挤压共同驱动,导致有效候选路径的采样概率彻底消失。尽管KL正则化试图缓解此问题,但其强加的全局形状匹配约束迫使策略完全模仿参考模型密度,与正确性所需的锐化产生梯度冲突。为此,我们提出锚定策略优化(APO),将范式从全局形状匹配转向支持集覆盖。通过定义基于参考模型高置信度支持集的安全流形,APO允许高效锐化,同时在纠错时选择性引入恢复力以防止崩溃。理论推导表明,APO是最大化支持覆盖的梯度对齐机制,支持弹性恢复,重新激活有效分支。在数学基准上的实证评估显示,APO打破准确率与多样性之间的权衡,在显著提升Pass@1的同时,恢复了标准策略梯度方法通常丢失的Pass@K多样性。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Rewards (RLVR) is increasingly viewed as a tree pruning mechanism. However, we identify a systemic pathology termed Recursive Space Contraction (RSC), an irreversible collapse driven by the combined dynamics of positive sharpening and negative squeezing, where the sampling probability of valid alternatives vanishes. While Kullback-Leibler (KL) regularization aims to mitigate this, it imposes a rigid Shape Matching constraint that forces the policy to mimic the reference model's full density, creating a gradient conflict with the sharpening required for correctness. We propose Anchored Policy Optimization (APO), shifting the paradigm from global Shape Matching to Support Coverage. By defining a Safe Manifold based on the reference model's high-confidence support, APO permits aggressive sharpening for efficiency while selectively invoking a restorative force during error correction to prevent collapse. We theoretically derive that APO serves as a gradient-aligned mechanism to maximize support coverage, enabling an Elastic Recovery that re-inflates valid branches. Empirical evaluations on mathematical benchmarks demonstrate that APO breaks the accuracy-diversity trade-off, significantly improving Pass@1 while restoring the Pass@K diversity typically lost by standard policy gradient methods.

强化学习策略优化探索效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。