arXiv:2609.06386cs.LGcs.CL2026-09

发现大模型生成答案的验证误差存在强相关,影响评估可靠性。

Are Verifier Errors Independent Within a GRPO Group? Evidence from Qwen2.5 Rollouts

  • 分析8个答案组内验证误差相关性,发现相关系数达0.53
  • 同一提示下答案格式相似者误差更集中,如分数、根式等
  • 建议评估时考虑提示难度与答案形式,而非仅看平均错误率

基于自动验证器的分组强化学习(RLVR)对每个提示生成多个完成结果并评分。本文在MATH、GSM8K和DeepMath-103K数据集上,分析由Qwen2.5-1.5B生成的24,998个八答案组的验证误差依赖性。估计组内验证误差的合并相关系数为0.530(95%置信区间:0.500–0.560),在交换对称误差模型下,每组有效样本量降至1.70。不同答案格式的误差聚集程度差异显著:分数、根式、符号表达式和区间比单位标注和百分号更集中。在四种规则验证器配置下重播组间优势,发现最多0.83%的组存在优势方向不一致。由于每组重复采样同一提示,组内聚集可能反映共有的提示难度或答案格式,本文未区分两者。与多评估者判断相关性的研究不同,本分析聚焦同一验证器对多个完成结果的评分依赖性。这些发现推动应考虑提示和答案形式的验证噪声分析,而非仅依赖整体错误率。

原文摘要 · Abstract (English)

Group-based reinforcement learning with verifiable rewards (RLVR) scoresmultiple completions per prompt using automatic verifiers. Analysesbased on independent verifier errors may overlook dependence associatedwith shared answer formats. We investigate this dependence in24,998 groups of eight completions generated by Qwen2.5-1.5B onMATH, GSM8K, and DeepMath-103K. We estimate a pooled within-groupverifier-error correlation of 0.530 (95% confidence interval:0.500--0.560). Under an exchangeable-error model, this correspondsto a design-effect-adjusted effective sample size of 1.70 for aneight-completion group. Dependence varies substantially across answerforms: fractions, radicals, symbolic expressions, and intervals exhibitstronger clustering than unit annotations and percent signs. Replayinggroup-relative advantages across four rule-based verifier configurationsidentifies at least one advantage-sign disagreement in up to 0.83% ofgroups. Because a group is repeated sampling for one prompt, thiswithin-group clustering may reflect shared prompt difficulty as well asshared answer form, and we do not attempt to separate the two here.Unlike studies of correlated judgments across multiple evaluators, ouranalysis examines dependence across completions scored by the sameverifier. These findings motivate prompt- and answer-form-aware analysesof verifier noise rather than characterizations based solely onaggregate error rates.

验证误差模型评估大模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。