arXiv:2506.13639cs.CL2025-06被引 26

研究大模型当裁判的可靠性,发现评判标准和采样方式影响最大。

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability

  • 用不同评判标准和采样策略测试大模型评估效果
  • 非确定性采样让机器评分更贴近人类偏好
  • 清晰标准下思维链对评分提升不大

随着大语言模型持续发展,可靠评估方法对开放性指令遵循任务尤为重要。LLM-as-a-Judge 通过大模型作为评估者实现自动化评估,但其可靠性仍存疑。本文针对与人类判断的一致性及评估一致性,分析了关键设计因素的影响。基于 BIGGENBench 与 EvalBiasBench,研究了评估设计、解码策略及思维链(CoT)推理在评估中的作用。结果表明:评估标准是影响可靠性的核心因素;非确定性采样相比确定性评估能更好对齐人类偏好;当存在明确评估标准时,思维链推理带来的提升有限。

原文摘要 · Abstract (English)

As large language models (LLMs) continue to advance, reliable evaluation methods are essential particularly for open-ended, instruction-following tasks. LLM-as-a-Judge enables automatic evaluation using LLMs as evaluators, but its reliability remains uncertain. In this work, we analyze key factors affecting its trustworthiness, focusing on alignment with human judgments and evaluation consistency. Using BIGGENBench and EvalBiasBench, we study the effects of evaluation design, decoding strategies, and Chain-of-Tought (CoT) reasoning in evaluation. Our results show that evaluation criteria are critical for reliability, non-deterministic sampling improves alignment with human preferences over deterministic evaluation, and CoT reasoning offers minimal gains when clear evaluation criteria are present.

大模型评估LLM裁判评估设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。