arXiv:2605.30504cs.CL2026-05中稿 · EMNLP被引 2

用心理测量学方法检测大模型评测中的错误标签。

Auditing LLM Benchmarks with Item Response Theory

论文配图:Auditing LLM Benchmarks with Item Response Theory
图 1 · 摘自论文原文
  • 基于项目反应理论识别评测数据中可能的错标项。
  • 在7个基准上以95%精度定位前200个错标样本。
  • 适合评估大模型评测数据质量的研究者使用。

大模型评测标签在发布后固定不变,错误被无声传播至下游评测。本文提出一种基于项目反应理论的指标,在114个模型的响应基础上,以95%精度在7个偏好和多选基准中识别出前200个潜在错标项,优于有监督分类器。错误源于机械标注规则、上游数据集继承的标注失误,以及本身无明确唯一答案的模糊题项。同一模型拟合还发现,奖励模型更擅长判断风格偏好而非事实知识;其中一前沿模型对检测出的错标项有78%准确率,远高于同行38%,表明存在评测污染或评测特定过度优化。

原文摘要 · Abstract (English)

LLM benchmark labels are frozen at release and silently propagated into downstream benchmarks, errors and all. We introduce an Item Response Theory-based indicator that surfaces likely mislabels at 95% precision in the top 200 examples across seven preference and multiple-choice benchmarks using responses from 114 models, outperforming a supervised classifier. We trace these errors to mechanical labeling heuristics, upstream annotation mistakes inherited unchanged from source datasets, and fundamentally ambiguous items without a defensible single label. The same model fit reveals that reward models specialize in stylistic preference rather than factual knowledge, and identifies one frontier reward model that agrees with detected mislabels at 78% accuracy versus 38% for its peers, consistent with benchmark contamination or benchmark-specific over-optimization.

大模型评测项目反应理论标签质量奖励模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。