构建首个指令遵循的细粒度评分基准,揭示大模型评分可靠性问题
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
- 设计首个基于评分标准的元评估基准,精细检验评分准确性
- 测试发现即使GPT-4o在难题集上准确率也仅55.97%
- 提出评分改进建议,适合评估系统设计者与评测研究人员
基于评分标准的评估已成为大语言模型指令遵循能力评价的主流方法。然而,这类评分的可靠性尚未明确,亟需元评估。现有元评估多聚焦于生成结果层面,未能检验评分标准所依赖的细粒度判断能力。为此,我们提出RubricEval基准:(1)首个针对指令遵循的评分层级元评估基准;(2)涵盖多种类别与模型来源的多样化指令与响应;(3)包含3,486个质量控制实例,以及易/难子集,更好区分评分员表现。实验显示,评分任务仍远未解决:即便广泛使用的GPT-4o在难题集上准确率仅为55.97%。分析表明,评分标准优于清单式评估,显式推理可提升准确率,并共同降低评分员间差异。基于建立的评分分类体系,我们识别出常见失败模式,为可靠评估提供可操作洞见。
原文摘要 · Abstract (English)
Rubric-based evaluation has become a prevailing paradigm for evaluating instruction following in large language models (LLMs). Despite its widespread use, the reliability of these rubric-level evaluations remains unclear, calling for meta-evaluation. However, prior meta-evaluation efforts largely focus on the response level, failing to assess the fine-grained judgment accuracy that rubric-based evaluation relies on. To bridge this gap, we introduce RubricEval. Our benchmark features: (1) the first rubric-level meta-evaluation benchmark for instruction following, (2) diverse instructions and responses spanning multiple categories and model sources, and (3) a substantial set of 3,486 quality-controlled instances, along with Easy/Hard subsets that better differentiates judge performance. Our experiments reveal that rubric-level judging remains far from solved: even GPT-4o, a widely adopted judge in instruction-following benchmarks, achieves only 55.97% on Hard subset. Considering evaluation paradigm, rubric-level evaluation outperforms checklist-level, explicit reasoning improves accuracy, and both together reduce inter-judge variance. Through our established rubric taxonomy, we further identify common failure modes and offer actionable insights for reliable instruction-following evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。