检验IRT在大模型评估中是否可靠,发现传统方法易失效,新方法有误判风险。
Can We Trust Item Response Theory for AI Evaluation?

- 用六大数据集模拟不同场景,测试四种IRT估计器表现
- 小样本或非正态分布下,排名和性能预测可能严重失真
- 提醒研究者注意样本量与诊断工具,避免误读评估结果
AI评测越来越多地使用项目反应理论(IRT)等项目级统计模型来估计模型能力、排序系统、筛选有信息量的题目并诊断评测质量。然而,AI评测数据常偏离人类测试的数据特征:模型数量少、题目数量多,且能力分布可能偏斜、聚集或呈多峰状。本文基于六个广泛使用的大型语言模型评测数据,提取项目参数和能力分布,模拟三种常见IRT模型下的答题矩阵,对比近期评测研究中使用的四种估计方法:边际最大似然、马尔可夫链蒙特卡洛、变分推断及神经伪西门子估计器。在18,000种模拟条件下,系统评估了计算可行性、可扩展性以及对模型排名、预测性能和题目特性的推断可靠性。结果表明,经典估计器在大规模评测中可能不可行,而可扩展估计器在小样本或非正态分布模型集合下会产生不可靠的项目级推断和排名结论。本研究明确了潜在特质模型在何种条件下能可信支持评测结论,以及实现可信使用所需的数据规模与诊断手段。
原文摘要 · Abstract (English)
AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT estimation tools were originally developed: benchmarks typically involve fewer evaluated models, far more items, and capability distributions that may be skewed, clustered, or multimodal. We examine how these regime mismatches challenge the reliability of IRT modeling for AI evaluation. Using item parameters and capability distributions derived from six widely used LLM benchmarks, we simulate response matrices under three common IRT models and compare four estimation tools used in recent benchmark studies: marginal maximum likelihood, Markov chain Monte Carlo, variational inference, and a neural pseudo-Siamese estimator. Across 18,000 simulation conditions, we systematically evaluate computational feasibility, scalability, and the reliability of IRT inferences about model rankings, predicted performance, and item characteristics. Results show that classical estimators can become infeasible in large benchmark settings, whereas scalable estimators can produce unreliable item-level and ranking inferences with small or non-normally distributed model sets. This study identifies when latent trait models reliably support or risk distorting AI benchmarking claims, and what sample sizes and diagnostics are needed for trustworthy use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。