arXiv:2606.00437cs.LG2026-06被引 1

提出测试语言模型过程奖励模型稳定性的新框架,发现现有模型在推理结构调整后评分易失真。

EST-PRM: Stress-Testing Process Reward Models Before They Become Load-Bearing

论文配图:EST-PRM: Stress-Testing Process Reward Models Before They Become Load-Bearing
图 1 · 摘自论文原文
  • 设计三种推理结构扰动方法,模拟真实场景中推理路径变化
  • 数学推理模型在位置调整下评分相关性下降0.152,分数虚高达32.8%
  • 揭示奖励模型对正确性判断的不一致,适合评估和优化训练机制

过程奖励模型(PRMs)广泛用于语言模型的密集步骤监督训练,其假设是经过标签保持变换后的推理步骤评分仍能稳定反映步骤正确性。然而,此类变换虽保留最终答案但改变推理结构,可能使评分与正确性信号的关系发生偏移,导致模型出现不同故障模式。为此,本文提出EST-PRM——一种针对密集过程奖励的应力测试框架,包含三种变换:(1)步骤膨胀,(2)依赖感知步骤重排,(3)置信度标记。定义了漏洞分解,将奖励虚高与正确性敏感性下降分离。在来自MATH-500、GSM8K和PRMBench的4,687条推理链上评估五种类PRM模型。结果表明各模型漏洞模式差异明显:Math-Shepherd对位置扰动最敏感,皮尔逊相关性下降0.152±0.038,评分虚高率32.8±4.9%;Qwen2.5-Math-PRM受步骤膨胀影响最大,虚高率高达47.6±4.3%。置信度扰动也导致奖励校准失真,暴露出正确性估计的不一致性。三种缓解策略被验证,揭示鲁棒性覆盖与误报率之间的权衡。

原文摘要 · Abstract (English)

Process reward models (PRMs) are widely used in language-model training with dense step-level supervision. They assume PRM scores are stable proxies for step correctness under label-preserving transformations. These transformations change reasoning structure but preserve final answers. We argue this assumption is not well validated. Such transformations can change how PRM scores relate to correctness signals, leading to different failure modes across models.To address this gap, we introduce \textbf{EST-PRM}, a stress-testing framework for dense process rewards. It applies three transformations: (1) step inflation, (2) dependency-aware step reordering, and (3) confidence markers. A vulnerability decomposition is defined that separates reward inflation from loss of correctness sensitivity. Five PRM-style models are evaluated on 4,687 reasoning chains from MATH-500, GSM8K, and PRMBench.The results indicate clear differences in vulnerability patterns across models. Math-Shepherd shows the strongest sensitivity to position perturbations, with a Pearson correlation drop of $0.152 \pm 0.038$ and a $32.8 \pm 4.9\%$ score inflation rate. Qwen2.5-Math-PRM is most affected by step inflation, reaching a $47.6 \pm 4.3\%$ inflation rate. Confidence-based perturbations also distort reward calibration, revealing inconsistencies in correctness estimation. Three mitigation strategies are evaluated, highlighting trade-offs between robustness coverage and false-positive rates.

奖励模型推理评估稳定性测试数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。