arXiv:2505.13346cs.CLcs.AI2025-05ACL被引 9

用强化学习训练更公平的推理型模型评判器。

J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization

  • 提出新算法,减少评判时的位置偏差影响。
  • 在推理评测中,性能超越GPT-4o和同类小模型。
  • 适合需要高可靠自动评估的科研与工程场景。

为应对大语言模型快速发展的需求,模型输出评估正从耗时的人工评估转向自动评估,即由大语言模型自身担任评判角色。当前的LLM作为评判者在对话质量等简单任务上表现良好,但在涉及复杂推理的内容评估中表现不佳。为此,本文探索使用强化学习训练评判模型。主要贡献包括:(1) 提出等初态组相对策略优化(EIS-GRPO)算法,使评判模型在复杂评估场景中对位置偏差更具鲁棒性;(2) 构建全新基准 ReasoningJudgeBench,覆盖以往工作未涵盖的多样化推理评估场景;(3) 训练出70亿参数的J4R评判模型,基于EIS-GRPO,其在JudgeBench和ReasoningJudgeBench上分别优于GPT-4o 6.7%、优于次优小模型9%,达到甚至超过更大规模的GRPO训练评判模型的表现。

原文摘要 · Abstract (English)

To keep pace with the increasing pace of large language models (LLM) development, model output evaluation has transitioned away from time-consuming human evaluation to automatic evaluation, where LLMs themselves are tasked with assessing and critiquing other model outputs. LLM-as-judge models are a class of generative evaluators that excel in evaluating relatively simple domains, like chat quality, but struggle in reasoning intensive domains where model responses contain more substantive and challenging content. To remedy existing judge shortcomings, we explore training judges with reinforcement learning (RL). We make three key contributions: (1) We propose the Equivalent Initial State Group Relative Policy Optimization (EIS-GRPO) algorithm, which allows us to train our judge to be robust to positional biases that arise in more complex evaluation settings. (2) We introduce ReasoningJudgeBench, a benchmark that evaluates judges in diverse reasoning settings not covered by prior work. (3) We train Judge for Reasoning (J4R), a 7B judge trained with EIS-GRPO that outperforms GPT-4o and the next best small judge by 6.7% and 9%, matching or exceeding the performance of larger GRPO-trained judges on both JudgeBench and ReasoningJudgeBench.

模型评判强化学习推理评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。