arXiv:2605.07046stat.MLcs.AI2026-05

用更稳定高效的方法评估大模型能力,看清题目难易与模型水平

An Interpretable and Scalable Framework for Evaluating Large Language Models

论文配图:An Interpretable and Scalable Framework for Evaluating Large Language Models
图 1 · 摘自论文原文
  • 基于极大化极小化原理重构评估问题,提升计算稳定性
  • 在MATH-500等数据集上速度提升数量级,精度相当或更高
  • 可揭示题目难度与区分度,适合优化评测体系的设计

大语言模型评估日益重要,但传统基准方法仅依赖平均准确率,忽视了模型输出的随机性及题目间的异质性。项目反应理论(IRT)虽能建模隐含的模型能力与题目特征,但传统方法计算成本高且数值不稳定,难以大规模应用。为此,我们提出一种可解释且可扩展的评估框架,基于极大化极小化原则,将问题转化为一系列约束矩阵分解子问题,实现参数估计的稳定高效,并具备可识别性与收敛性理论保证。在合成数据和真实数据集(包括MATH-500及六个Open LLM Leaderboard基准)上的实验表明,该方法在保持或优于现有方法精度的同时,相较竞品实现数量级的速度提升。结果符合已知缩放规律,可揭示题目难度与区分度,为更科学的评测设计提供支持。

原文摘要 · Abstract (English)

Evaluation of large language models (LLMs) is increasingly critical, yet standard benchmarking methods rely on average accuracy, overlooking both the inherent stochasticity of LLM outputs and the heterogeneity of benchmark items. Item Response Theory (IRT) offers a principled framework for modeling latent model abilities and item characteristics, but conventional methods are computationally expensive and numerically unstable, limiting large-scale implementations. To address these challenges, we propose an interpretable and scalable framework for LLM evaluation based on the majorization-minimization principle. Our approach reformulates the problem as a sequence of constrained matrix factorization subproblems, enabling stable and efficient parameter estimation with theoretical guarantees for identifiability and convergence. Experiments on synthetic and real-world datasets, including MATH-500 and six Open LLM Leaderboard benchmarks, demonstrate that our method achieves superior scalability and interpretability. It delivers orders-of-magnitude speedups over competing methods while maintaining comparable or even higher estimation accuracy. Our results align with established scaling laws and offer insights into item difficulty and discrimination, informing more principled benchmark design.

大模型评估项目反应理论可解释性高效计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。