arXiv:2607.25257stat.APcs.AI2026-07

给LLM评测的神经项目反应模型加了不确定性量化,让结果更可信。

Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks

论文配图:Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks
图 1 · 摘自论文原文
  • 用后验拉普拉斯近似,在不重训练的情况下添加贝叶斯不确定性估计。
  • 12个模型的对比中,多数差异不显著,点估计排名不可靠。
  • 新方法能稳定评估题目难度,适合做小样本评测和统计推断。

项目反应理论(IRT)最近被用于评估大语言模型(LLM)基准,将模型隐性能力与题目特性分离。现有神经IRT方法(如PSN-IRT)使用点估计,限制了不确定性量化与下游统计推断。本文提出Laplace-PSN-IRT,一种后处理的最后层拉普拉斯近似,为已训练的PSN-IRT模型增加近似贝叶斯后验推断,无需重训练即可恢复模型能力与题目难度的校准不确定性。该后验支持可信区间、模型间的概率比较,并可将参数不确定性传播至基于费雪信息的题目选择中。我们发现,在标准LLM基准排行榜上,12个模型的大多数成对比较在点估计排名差异下并不具备统计显著性。此外,点估计的费雪信息在多数题目上几乎为零,因其仅在单一参考能力下评估;而后验期望费雪信息在整个能力范围内保持更稳定。最终,后验期望费雪信息在多数实验设置中能更准确地从少量题目子集恢复全基准的能力排序,且在最小子集上性能与点估计相当。通过保留预测覆盖率验证了近似后验的校准性,结果显示在该架构中将题目难度视为随机而题目区分度固定,能产生校准良好的不确定性。

原文摘要 · Abstract (English)

Item Response Theory (IRT) has recently been proposed as a framework for evaluating large language model (LLM) benchmarks by separating a model's latent ability from the properties of individual benchmark items. Existing neural IRT approaches, including PSN-IRT, estimate these quantities using point estimates, limiting uncertainty quantification and downstream statistical inference. We introduce Laplace-PSN-IRT, a post-hoc last-layer Laplace approximation that augments a trained PSN-IRT model with approximate Bayesian posterior inference, recovering calibrated uncertainty over model ability and item difficulty without retraining. The resulting posterior enables credible intervals, probabilistic comparisons between models, and propagation of parameter uncertainty into Fisher-information-based item selection. We show that most pairwise comparisons among 12 models on a standard LLM benchmark leaderboard are not statistically distinguishable despite differing point-estimate ranks. We further show that point-estimate Fisher information can become nearly zero for many benchmark items because it is evaluated at a single reference ability, whereas posterior-expected Fisher information remains substantially more stable across the ability range. Finally, posterior-expected Fisher information more accurately recovers full-benchmark ability rankings from small benchmark subsets in most experimental settings while matching point-estimate performance for the smallest subsets. We validate the calibration of the approximate posterior using held-out predictive coverage and find that modeling item difficulty as random while treating item discrimination as fixed produces well-calibrated uncertainty in this architecture.

LLM评估不确定性量化项目反应理论贝叶斯推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。