RLVR模型性能评估存在三重偏差,需匹配预算与数据防污染才可信。
Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards
- 通过匹配预算和数据版本复现实验,消除评估偏差
- 部分公开结果因数据污染导致性能虚高,真实提升有限
- 提出含校准与拒答追踪的标准化评估框架,适合部署前验证
基于可验证奖励的强化学习(RLVR)是提升大语言模型在数学、代码等结构化任务上的有效方法。然而,我们指出当前许多宣称的性能提升尚未充分验证,因其混杂了三种干扰因素:(i)RLVR与基线评估间预算不匹配;(ii)拒绝回答被误判为自信输出的尝试膨胀与校准漂移;(iii)基准数据集污染。通过预算匹配的复现和部分提示污染探测,我们发现多个广受引用的性能差距在统一预算、提示和数据版本后显著缩小甚至消失。这并非否定RLVR的有效性,而是表明现有测量常夸大能力提升并掩盖可靠性代价。因此,我们提出一套简洁、税感明确的最低标准:预算匹配的饱和曲线(带方差)、校准与拒答追踪、使用大模型裁判时的鲁棒性压力测试,以及显式污染筛查。在此控制下,RLVR仍具有效性且可部署,但推理提升必须在无此控制时视为暂定结论。
原文摘要 · Abstract (English)
Reinforcement learning with verifiable rewards (RLVR) is a practical, scalable way to improve large language models on math, code, and other structured tasks. However, we argue that many headline RLVR gains are not yet well validated because reports often conflate policy improvement with three confounds: (i) budget mismatch between RLVR and baseline evaluations, (ii) attempt inflation and calibration drift that convert abstentions into confident answers, and (iii) benchmark data contamination. Using budget-matched reproductions and partial-prompt contamination probes, we find that several widely cited gaps shrink substantially or disappear once budgets, prompts, and dataset versions are matched and contaminated sets are treated as memorization probes rather than evidence of reasoning. This does not mean that RLVR is ineffective, but it implies that current measurements often overstate capability gains and obscure reliability costs. We therefore propose a compact, tax-aware minimum standard for RLVR training and evaluation: budget-matched saturation curves with variance, calibration, and abstention tracking, a judge-robustness stress test when LLM judges are used, and an explicit contamination screen. With these controls, RLVR remains effective and deployable in verifiable domains, but reasoning gains should be treated as provisional without them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。