arXiv:2607.20543cs.LGcs.AI2026-07

RLVR训练后反而降低多轮采样成功率,本文揭示其机制并提出改进方法。

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion

  • 发现RLVR会丢失基础模型在边界提示下的稀有正确路径
  • 提出PBA方法通过锚定基线分布,提升高预算覆盖率
  • 适用于需反复验证的推理系统优化,如数学与视觉语言模型

基于可验证奖励的强化学习(RLVR)虽能提升单次采样准确率,但在多次采样时可能使模型表现更差。本文研究这种pass@k反转现象:训练后策略解决的独立问题数反而少于基线模型。失败集中在边界提示上,因基线模型存在稀有的正确轨迹,虽可通过采样恢复但难以在有限的RLVR回放组中稳定出现。本文提出二模解释:这是‘无证据’失效——稀有正确路径在采样前消失,而RLVR未足够强化它们。核心贡献是诊断与机制分析。提出的每问题基线锚定(PBA)是一种简单验证:用冻结基线的正确证据锐化提示,并将高风险提示锚定至基线分布。在Omni-MATH-Test三个训练种子上,以MATH500为高覆盖验证基准,PBA同时优于匹配的GRPO在Pass@1和高预算覆盖率。3000提示的受控诊断实验在各种子间一致显示:普通GRPO丢失基线可解的边界提示,而PBA保留了稀有验证通过的轨迹。本文使用数学验证器作为可控测试床,验证引导优化;同样的pass@k反转风险也存在于ECCV相关的视觉语言代理中,当重复的视觉、空间或图表推理被外部工具验证时。推理后训练应权衡优化强度与安全提示的选择。

原文摘要 · Abstract (English)

Reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while making a model worse under repeated sampling. We study this pass@k inversion: after training, the policy may solve fewer distinct problems than its base model at large $k$. The failure concentrates on boundary prompts, where the base model contains rare correct trajectories that are recoverable by sampling but too sparse to reliably appear in finite RLVR rollout groups. We argue that a two-mode account explains this as an absence-of-evidence failure: rare correct trajectories may disappear before RLVR samples and reinforces them often enough. The main contribution is this diagnostic and mechanistic framing. Per-Problem Base Anchoring (PBA) is a deliberately simple proof-of-concept: sharpen prompts with sufficient frozen-base correct evidence, and anchor risky prompts to the base distribution. Across three training seeds on Omni-MATH-Test, with MATH500 as a secondary high-coverage validation benchmark, PBA improves both \PassK{1} and high-budget coverage over matched GRPO. A 3000-prompt regime-controlled diagnostic study is consistent across seeds with the expected signature: ordinary GRPO loses base-solvable boundary prompts, while PBA preserves rare verifier-positive trajectories. We use mathematical verifiers as a controlled testbed for verifier-guided optimization; the same pass@k inversion risk applies to ECCV-relevant vision-language agents when repeated visual, spatial, or chart-reasoning attempts are checked by external tools or verifiers. Reasoning post-training should decide not only how strongly to optimize, but which prompts are safe to optimize.

强化学习推理优化数学推理验证器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。