arXiv:2602.22585cs.AIcs.LG2026-02被引 1

用心理测量模型修正人工评分偏差,提升AI评估可靠性

Correcting Human Labels for Rater Effects in AI Evaluation: An Item Response Theory Approach

  • 引入项目反应理论建模评分者偏差,分离真实质量与评分行为
  • 在OpenAI摘要数据集上修正后,评分一致性显著提升
  • 适合关注AI评估公平性与可解释性的研究者和开发者

人工评估在训练和评测AI模型中起核心作用,但这些数据常被当作无系统误差的测量值。本文将心理测量评分者模型融入AI流程,提升基于人工判断结论的可靠性和有效性。文章回顾了常见的评分者效应——严厉度与中心化倾向,并展示如何利用多面拉斯克模型等项目反应理论方法,分离出真实输出质量与评分者行为。以OpenAI摘要数据集为例,调整评分者严厉度后,得到更准确的摘要质量估计,并可诊断评分者表现。将心理测量建模纳入人机协作评估,使人工数据使用更严谨透明,支持开发者基于校正后的分数做决策而非原始有偏评分。这为构建更稳健、可解释且符合构念的AI开发与评估实践指明方向。

原文摘要 · Abstract (English)

Human evaluations play a central role in training and assessing AI models, yet these data are rarely treated as measurements subject to systematic error. This paper integrates psychometric rater models into the AI pipeline to improve the reliability and validity of conclusions drawn from human judgments. The paper reviews common rater effects, severity and centrality, that distort observed ratings, and demonstrates how item response theory rater models, particularly the multi-faceted Rasch model, can separate true output quality from rater behavior. Using the OpenAI summarization dataset as an empirical example, we show how adjusting for rater severity produces corrected estimates of summary quality and provides diagnostic insight into rater performance. Incorporating psychometric modeling into human-in-the-loop evaluation offers more principled and transparent use of human data, enabling developers to make decisions based on adjusted scores rather than raw, error-prone ratings. This perspective highlights a path toward more robust, interpretable, and construct-aligned practices for AI development and evaluation.

AI评估心理测量评分者偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。