给评分模型提供额外线索,让普通模型也能精准评估顶尖大模型。
Graders should cheat: privileged information enables expert-level automated evaluations
- 用真实答案或解题指引作为额外信息,提升评分模型能力。
- 在奥数等难题上,评分模型表现接近人类专家,超越多数专用模型。
- 适合想低成本评估大模型性能的研究者与工程师使用。
自动评估语言模型(即用一个评分模型评判候选模型)能显著加速并降低评估成本。但存在悖论:若评分模型本就弱于候选模型,如何可靠评估超出两者能力边界的问题?例如当前模型在研究生级物理和奥数级数学上表现不佳,难以胜任评分任务。本文表明,提供特权信息(如标准答案或问题专用指导)可显著提升对前沿难题的自动化评估效果。该方法带来两大优势:一是扩大了评分模型的应用范围,使较弱模型能评价更强模型;二是利用特权信息设计更简单的变体问题,增强不同模型在低性能任务上的区分度。实验显示,通用评分模型在RewardBench上达到最先进水平,超过几乎所有专门调优的模型;在Vibe-Eval上优于单个真人评分员,在奥数级数学题上接近人类专家表现。
原文摘要 · Abstract (English)
Auto-evaluating language models (LMs), i.e., using a grader LM to evaluate the candidate LM, is an appealing way to accelerate the evaluation process and the cost associated with it. But this presents a paradox: how can we trust the grader LM, which is presumably weaker than the candidate LM, to assess problems that are beyond the frontier of the capabilities of either model or both? For instance, today's LMs struggle on graduate-level physics and Olympiad-level math, making them unreliable graders in these domains. We show that providing privileged information -- such as ground-truth solutions or problem-specific guidelines -- improves automated evaluations on such frontier problems. This approach offers two key advantages. First, it expands the range of problems where LMs graders apply. Specifically, weaker models can now rate the predictions of stronger models. Second, privileged information can be used to devise easier variations of challenging problems which improves the separability of different LMs on tasks where their performance is generally low. With this approach, general-purpose LM graders match the state of the art performance on RewardBench, surpassing almost all the specially-tuned models. LM graders also outperform individual human raters on Vibe-Eval, and approach human expert graders on Olympiad-level math problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。