用心理测量学方法评估大模型评分能力,发现其与人类差异显著。
Rating the Raters: Rasch Measurement Theory for LLM Evaluation

- 基于拉什模型拆分评分中的多个影响因素
- 九个大模型评分存在严重度、敏感性等系统偏差
- 适合研究模型评价机制或评测设计的学者
大语言模型如今既作为被测对象参与基准测试,也充当其他模型输出或人类内容的评分者。这些场景均可视为测量问题:通过工具(如测评集)对对象的潜在属性进行探测。当前标准评估常忽视各核心组件对结果的影响,限制了我们对测量本质的理解。拉什测量理论(RMT)正适合此类问题:它可将有序评分分解为可分离的维度,并在统一量表上实现。同时提供一系列诊断工具,识别测量失准和评分者偏差。本研究以仇恨言论评测语料库为案例,应用多面拉什模型分析九个来自不同家族和能力层级的大模型评分数据。结果显示,大模型在评分严重度、题目校准、题序鲁棒性、目标身份敏感性及评分尺度使用等方面系统偏离人类评分者,而这些差异在常规评估中会被掩盖。总体而言,我们主张将RMT纳入评估大模型作为被试、评判者和评分者的工具箱。
原文摘要 · Abstract (English)
LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models' outputs, and raters of human-generated content. Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an instrument (e.g., benchmark) by raters. Standard evaluation practices often neglect the contributions of each core component to the end result, limiting our understanding of what is being measured. Rasch measurement theory (RMT) is well-suited to this kind of problem. RMT decomposes ordinal ratings into separable facets on a common scale. It further provides a battery of diagnostics that can identify miscalibrated measurements and rater biases. We present a case study of RMT applied to the LLM-as-rater paradigm using the Measuring Hate Speech corpus, whose construct was itself built under RMT. We fit a series of many-facet Rasch models to annotations from nine LLMs spanning families and capability levels. Our analyses show that LLMs systematically differ from human raters in severity, item-level calibration, question-order robustness, target-identity sensitivity, and rating scale use, which all would be obscured by standard evaluation practice. Overall, we argue that RMT belongs in the toolkit for evaluating LLM-as-examinee, -judge, and -rater paradigms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。