arXiv:2503.05965cs.LGcs.CY2025-03NeurIPS被引 24

提出新方法解决大模型评分时因评判标准模糊导致的偏差问题。

Validating LLM-as-a-Judge Systems under Rating Indeterminacy

  • 引入多标签评分集应对评分不确定性,避免强制单选带来的偏差。
  • 实验证明传统方法选出的评分模型性能比新方法差最多31%。
  • 适合关注AI评估可靠性的研究者与工程师参考使用。

LLM-as-a-judge范式通过大模型替代人工评分,实现生成式AI评估的规模化与标准化。然而,许多评分任务存在评分标准多重合理解释的情况(称作评分不确定性),而当前普遍采用强制单选的评分方式,导致人类与大模型在处理不确定性时产生偏差。本文建立理论框架,分析不同评分收集与聚合方式对评价结果的影响。基于11个真实世界评分任务和9个商用大模型的实验表明,依赖强制单选的主流验证方法会选出性能显著更差的评分系统,最差情况下表现落后达31%。所提方法采用多标签评分集,有效缓解了该偏差,为更可靠的评估提供了可操作方案。

原文摘要 · Abstract (English)

The LLM-as-a-judge paradigm, in which a judge LLM system replaces human raters in rating the outputs of other generative AI (GenAI) systems, plays a critical role in scaling and standardizing GenAI evaluations. To validate such judge systems, evaluators assess human--judge agreement by first collecting multiple human ratings for each item in a validation corpus, then aggregating the ratings into a single, per-item gold label rating. For many items, however, rating criteria may admit multiple valid interpretations, so a human or LLM rater may deem multiple ratings "reasonable" or "correct." We call this condition rating indeterminacy. Problematically, many rating tasks that contain rating indeterminacy rely on forced-choice elicitation, whereby raters are instructed to select only one rating for each item. In this paper, we introduce a framework for validating LLM-as-a-judge systems under rating indeterminacy. We draw theoretical connections between different measures of judge system performance under different human--judge agreement metrics, and different rating elicitation and aggregation schemes. We demonstrate that differences in how humans and LLMs resolve rating indeterminacy when responding to forced-choice rating instructions can heavily bias LLM-as-a-judge validation. Through extensive experiments involving 11 real-world rating tasks and 9 commercial LLMs, we show that standard validation approaches that rely upon forced-choice ratings select judge systems that are highly suboptimal, performing as much as 31% worse than judge systems selected by our approach that uses multi-label "response set" ratings to account for rating indeterminacy. We conclude with concrete recommendations for more principled approaches to LLM-as-a-judge validation.

大模型评估评分系统可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。