arXiv:2607.12885cs.CL2026-07被引 1

LLM裁判在无标准答案时容易误判,加参考答案可让判断结果翻转85%。

LLM Judges Can Be Too Generous When There Is No Reference Answer

论文配图:LLM Judges Can Be Too Generous When There Is No Reference Answer
图 1 · 摘自论文原文
  • 通过校准与敏感性实验,测试LLM裁判在无参考答案下的判断能力。
  • 无参考答案时,错误回答被过度认可,正确率下降最高达85%。
  • 研究提供校准LLM裁判的实用方法,适合评估模型输出的研究者使用。

LLM裁判正被广泛用于评估开放式模型输出,常在无参考答案的场景中使用。然而,这种评估方式是否可靠?本文通过两阶段流程:一是校准实验,评估裁判模型对任务的理解;二是敏感性实验,分析参考答案在提示中的存在与位置对裁判表现的影响。覆盖三种语言的实验表明,缺乏参考答案时,裁判模型倾向于过度认可错误回答,添加参考答案信息可使正确/错误判断的决策翻转高达85%。与部分人工标注对比显示,这种参考驱动的变化通常与人类判断一致。结果强调,必须在有参考答案的样本上校准LLM裁判后,才能在无参考设置中可靠使用。本研究为其他任务中校准LLM裁判提供了可复用的方法论。

原文摘要 · Abstract (English)

LLM judges are increasingly being used to evaluate open-ended model responses, often in no-reference settings where a ground-truth answer is unavailable. However, can they reliably assess in such evaluation setups? We explore this question in this paper through a two stage pipeline with a) calibration experiments that assess the judge model's knowledge of the task it is evaluating, and b) sensitivity experiments that assess how the judge model's performance is impacted by the presence and positioning of the reference answer in the prompt. Across experiments covering three languages, we show that the judge models we evaluated tend to over-credit incorrect answers in the absence of a reference answer, and adding reference answer information to the prompt flips the judge model's correct/incorrect decisions by as much as 85% in some experimental settings. Comparison with a subset of human annotations shows that these reference-driven changes generally align with human judgments. Our results emphasize the need for calibrating the LLM judges with a sample with reference-aware evaluation before using them in reference-free setups reliably, and our methodology provides a blueprint for researchers and practitioners in doing such calibration of LLM judges for other tasks.

大模型评估LLM裁判无参考评价校准实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。