arXiv:2505.15055cs.CL2025-05AAAI被引 39

用心理测量学方法重评大模型评测,发现现有榜单有偏差,可构建更精准的小型评测集。

Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory

  • 引入改进的项目反应理论框架,同时评估题目难度与模型能力。
  • 分析4万多个题目发现主流评测存在测量质量差异。
  • 新方法可缩小评测规模但更贴近人类偏好,适合高效评估模型。

大型语言模型(LLM)的评测广泛依赖基准测试,但不同排行榜结果不一致、顶尖模型区分度差,暴露出其难以真实反映模型能力的问题。本文对主流大模型基准进行批判性分析,基于多样模型结果展开研究。提出一种基于项目反应理论的伪孪生网络(PSN-IRT),在经典IRT架构中引入丰富题项参数,实现对题项特性与模型能力的精准可靠估计。基于此框架,我们对包含41,871个题目的11个主流大模型基准进行了系统分析,揭示了其在测量质量上存在显著且多样的缺陷。此外,实验表明,利用PSN-IRT可构建更小规模的基准,仍能保持与人类偏好更强的一致性。

原文摘要 · Abstract (English)

The evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concerns about their ability to accurately reflect authentic model capabilities. This paper provides a critical analysis of benchmark effectiveness, examining mainstream prominent LLM benchmarks using results from diverse models. We first propose Pseudo-Siamese Network for Item Response Theory (PSN-IRT), an enhanced Item Response Theory framework that incorporates a rich set of item parameters within an IRT-grounded architecture. PSN-IRT can be utilized for accurate and reliable estimations of item characteristics and model abilities. Based on PSN-IRT, we conduct extensive analysis on 11 LLM benchmarks comprising 41,871 items, revealing significant and varied shortcomings in their measurement quality. Furthermore, we demonstrate that leveraging PSN-IRT is able to construct smaller benchmarks while maintaining stronger alignment with human preference.

大模型评测项目反应理论基准优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。