首次系统评估大模型裁判在模型排名中的表现与偏差。
JuStRank: Benchmarking LLM Judges for System Ranking
- 用多输出聚合评分构建模型系统排名,检验大模型裁判能力
- 发现裁判存在对特定系统的正负偏好,影响排名准确性
- 适合关注大模型评估公平性与可信赖性的研究者参考
随着生成式AI的快速发展,如何系统性地比较和选择众多模型与配置成为迫切需求。基于大模型的评判者因其规模和通用性,成为解决该问题的有力方案。关键在于需先验证评判者自身的质量。以往工作聚焦于对单个响应或响应对的评估,忽略了评判者对不同系统可能存在的偏见。本文首次开展大规模研究,将大模型作为系统排名的评判者。通过聚合多个系统输出的判断得分生成系统评分,并与人工排名对比以评估评判者质量。分析揭示了评判者的决策强度与系统偏好,为模型评估的可靠性提供了重要洞察。
原文摘要 · Abstract (English)
Given the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations available. The scale and versatility of such evaluations make the use of LLM-based judges a compelling solution for this challenge. Crucially, this approach requires first to validate the quality of the LLM judge itself. Previous work has focused on instance-based assessment of LLM judges, where a judge is evaluated over a set of responses, or response pairs, while being agnostic to their source systems. We argue that this setting overlooks critical factors affecting system-level ranking, such as a judge's positive or negative bias towards certain systems. To address this gap, we conduct the first large-scale study of LLM judges as system rankers. System scores are generated by aggregating judgment scores over multiple system outputs, and the judge's quality is assessed by comparing the resulting system ranking to a human-based ranking. Beyond overall judge assessment, our analysis provides a fine-grained characterization of judge behavior, including their decisiveness and bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。