arXiv:2509.22099cs.CL2025-09被引 4

解决生成模型在判断任务中的能力短板,提升评估准确性。

S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models

  • 用解题与判断并行监督,打通生成模型的求解与评估能力
  • 使判断错误率降低16.2%,判断性能提升5.8%
  • 无需外部强模型,小数据下实现自进化,适合高效训练GRM

随着大语言模型的快速发展,生成式奖励模型(GRMs)被广泛用于奖励建模与评估。以往研究主要通过偏好数据集优化专门的GRMs,以判断正确性为监督信号。尽管普遍认为更强的求解能力通常带来更好的判断能力,但我们首次在个体查询层面发现显著的‘求解-判断差距’:即便模型能正确求解某些问题,其判断正确率仍低至14%~37%。本文提出S2J方法,通过在同一模型输出上同时利用求解与判断能力作为监督信号,显式关联求解与评估能力,从而缩小该差距。大量实验表明,S2J将求解-判断差距降低16.2%,判断性能提升5.8%。值得注意的是,该方法在相同基座模型上达到当前最优表现,且训练数据量更少。更重要的是,S2J通过自演化实现,无需依赖更强大的外部模型进行知识蒸馏。

原文摘要 · Abstract (English)

With the rapid development of large language models (LLMs), generative reward models (GRMs) have been widely adopted for reward modeling and evaluation. Previous studies have primarily focused on training specialized GRMs by optimizing them on preference datasets with the judgment correctness as supervision. While it's widely accepted that GRMs with stronger problem-solving capabilities typically exhibit superior judgment abilities, we first identify a significant solve-to-judge gap when examining individual queries. Specifically, the solve-to-judge gap refers to the phenomenon where GRMs struggle to make correct judgments on some queries (14%-37%), despite being fully capable of solving them. In this paper, we propose the Solve-to-Judge (S2J) approach to address this problem. Specifically, S2J simultaneously leverages both the solving and judging capabilities on a single GRM's output for supervision, explicitly linking the GRM's problem-solving and evaluation abilities during model optimization, thereby narrowing the gap. Our comprehensive experiments demonstrate that S2J effectively reduces the solve-to-judge gap by 16.2%, thereby enhancing the model's judgment performance by 5.8%. Notably, S2J achieves state-of-the-art (SOTA) performance among GRMs built on the same base model while utilizing a significantly smaller training dataset. Moreover, S2J accomplishes this through self-evolution without relying on more powerful external models for distillation.

生成模型奖励模型自进化评估优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。