arXiv:2605.25252cs.LGcs.AI2026-05被引 1

实验证明:强化学习中纠错器质量比算力更重要。

Quantifying Empirical Compute-Supervision Tradeoffs in RLVR

论文配图:Quantifying Empirical Compute-Supervision Tradeoffs in RLVR
图 1 · 摘自论文原文
  • 用噪声干扰正确性信号,测试算力与监督质量的权衡
  • 算力增加无法弥补错误标签导致的性能差距,收益递减
  • 误判为错误比误判为正确危害更大,应优先降低假阴性

基于可验证奖励的强化学习(RLVR)已成为语言模型后训练的标准范式,但实际中验证器很少完美。近期理论预测表明,验证器噪声会影响学习速度但不影响最终结果,意味着充分算力可弥合不完美监督带来的差距。我们通过在GSM8K数据集上对Qwen2.5(0.5B、1.5B)使用GRPO进行后训练,向二元正确性信号注入可控的假阳性与假阴性噪声,并以每提示词的采样轨迹数作为算力指标进行测试。结果显示,在大幅增加算力的情况下,验证精度差距依然存在,且算力回报急剧递减。进一步发现,假阴性对性能的负面影响远高于假阳性,呈现结构性不对称。这表明验证器质量与训练算力不可互换,降低假阴性比单纯扩大算力更有效。

原文摘要 · Abstract (English)

Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training language models, but in practice, verifiers are rarely perfect. Recent theoretical work predicts that verifier noise affects the rate of learning but not its final outcome, implying that sufficient compute should close any gap induced by imperfect supervision. We test this prediction empirically by post-training Qwen2.5 (0.5B, 1.5B) with GRPO on GSM8K while injecting controlled false-positive and false-negative noise into the binary correctness signal, and varying rollouts per prompt as a compute axis. In practice, the gap in validation accuracy persists under substantial compute scaling, with returns to compute that are sharply diminishing. We further find a structural asymmetry where false negatives monotonically degrade performance more quickly than false positives. These findings suggest verifier quality and training compute are not interchangeable, and that reducing false negatives is a more effective lever than scaling compute alone.

强化学习语言模型验证器训练效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。