用心理测量学方法评估大模型评分可靠性,找出其不一致根源。
Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory
- 基于项目反应理论,分析大模型在不同提示下的评分稳定性
- 发现大模型评分一致性与人类评估存在显著偏差
- 为提升自动化评测可信度提供可解释的诊断工具
尽管大模型作为评分者被广泛用于自动化评估,现有验证方法主要关注输出结果,难以判断大模型评分是否具备稳定可靠的测量能力。为此,我们提出一个基于项目反应理论(IRT)的两阶段诊断框架,采用分级反应模型(GRM),从两个互补维度衡量可靠性:(1) 内在一致性,即在提示变化下评分行为的稳定性;(2) 人类对齐度,反映与人类质量评估的一致性。我们通过该框架实证检验了多种大模型评分者,结果表明,使用IRT-GRM能系统性地生成可解释的诊断信号,为验证大模型评分可靠性及识别不可靠原因提供实用指导。
原文摘要 · Abstract (English)
While LLM-as-a-Judge is widely used in automated evaluation, existing validation practices primarily operate at the level of observed outputs, offering limited insight into whether LLM judges themselves function as stable and reliable measurement instruments. To address this limitation, we introduce a two-phase diagnostic framework for assessing reliability of LLM-as-a-Judge, grounded in Item Response Theory (IRT). The framework adopts Graded Response Model (GRM) of IRT and formalizes reliability along two complementary dimensions: (1) intrinsic consistency, defined as the stability of measurement behavior under prompt variations, and (2) human alignment, capturing correspondence with human quality assessments. We empirically examine diverse LLM judges with this framework, and show that leveraging IRT-GRM yields interpretable signals for diagnosing judgments systematically. These signals provide practical guidance for verifying reliablity of LLM-as-a-Judge and identifying potential causes of unreliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。