对比两种评估方法,发现偏好判断更准且省时。
JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment

- 用配对偏好替代评分表,让专家更快更准评价输出质量
- 偏好判断的排序相关性是评分表的6倍以上(0.908 vs. 0.150)
- 适合需要专家评估但无标准答案的高专业领域研究
当前评估方法主要有两类:基于评分标准的打分和成对偏好判断。尽管广泛应用,二者选择缺乏依据。我们发布JudgmentBench,包含30个真实法律任务,配套1,539条评分标准打分和1,530组成对偏好判断,均由经验丰富的执业律师(来自美国顶级律所)完成。这是首个在高专业领域中,同一专家对相同任务同时提供两种标注信号的公开数据集。使用三种不同质量层级的LLM生成输出进行初步实证比较:在每任务秩相关性指标上,偏好判断的平均斯皮尔曼相关系数达0.908,远高于评分表的0.150(估计差异0.758 [0.494, 1.021]);在每判断胜负率上,偏好判断胜率0.669,优于评分表的0.542(估计差异0.127 [0.067, 0.186]),且标注时间不足一半。该模式在人类与LLM自动评分中均成立。数据集的配对结构为专家判断的采集、聚合及作为无真值领域监督信号的研究提供了新方向。
原文摘要 · Abstract (English)
Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative judgment elicits pairwise preferences between outputs. Although both methodologies are widely used, the choice between them is rarely justified. We release JudgmentBench, a benchmark of 30 real-world legal tasks, paired with 1,539 rubric scores and 1,530 pairwise preference judgments collected from practicing attorneys--including at major U.S. law firms--with substantial experience. The annotations constitute the first publicly available dataset in a high-expertise domain in which both supervision signals are elicited from the same experts on the same items. Using LLM-generated outputs at three constructed quality levels, we provide an initial empirical comparison: comparative judgments recover the intended quality ordering substantially better than rubrics under both a per-task rank-correlation metric (mean Spearman's rank correlation of 0.908 vs. 0.150, estimated difference = 0.758 [0.494, 1.021]) and a per-judgment pairwise win-rate metric (0.669 vs. 0.542, estimated difference = 0.127 [0.067, 0.186]), while requiring less than half the annotation time. The patterns hold for human annotators and LLM autograders. Beyond this initial comparison, the paired structure of the dataset supports a broader research agenda on how expert judgment should be elicited, aggregated, and used as supervision in domains without verifiable ground truth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。