arXiv:2512.07019stat.MEcs.AI2025-12被引 7

用新模型同时评估大模型回答准确率和推理长度,更全面反映其真实能力。

Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length

  • 融合回答准确率与思维链长度,引入能力-速度关联参数建模
  • 相比传统方法,估计更准、置信区间更短,预测力更强
  • 适合需要精细评估大模型推理能力的研究者和开发者

大型语言模型(LLMs)的快速发展迫切需要有效的评估方法以指导下游应用和未来改进。项目反应理论(IRT)结合计算机自适应测试已成为评估LLM回答准确率的有力框架。然而,仅关注准确率不足以反映模型的推理能力,其思维链(CoT)长度是重要指标。为此,我们提出新的延迟-响应理论(LaRT)模型,通过引入潜在能力与潜在速度间的相关性参数,联合建模响应准确率与CoT长度。我们推导出高效的随机近似期望最大化算法进行参数估计,并建立了潜在能力与速度参数的可识别性理论,保障估计的统计有效性。理论渐近分析与模拟研究均表明,LaRT在潜质估计精度和置信区间长度上优于IRT。在真实数据上,我们在多个主流基准数据集上收集了不同LLMs的响应结果,发现LaRT给出的模型排名不同于IRT,且在预测能力、题目效率、排序有效性及评估效率等多项关键指标上表现更优。代码与数据已公开于https://github.com/Toby-X/Latency-Response-Theory-Model。

原文摘要 · Abstract (English)

The proliferation of Large Language Models (LLMs) necessitates valid evaluation methods to provide guidance for both downstream applications and actionable future improvements. The Item Response Theory (IRT) model with Computerized Adaptive Testing has recently emerged as a promising framework for evaluating LLMs via their response accuracy. Beyond simple response accuracy, LLMs' chain of thought (CoT) lengths serve as a vital indicator of their reasoning ability. To leverage the CoT length information to assist LLM evaluation, we propose the \textbf{La}tency-\textbf{R}esponse \textbf{T}heory (LaRT) model, which jointly models both the response accuracy and CoT length by introducing a key correlation parameter between the latent ability and the latent speed. We derive an efficient stochastic approximation Expectation-Maximization algorithm for parameter estimation. We establish rigorous identifiability results for the latent ability and latent speed parameters to ensure the statistical validity of their estimation. Through both theoretical asymptotic analyses and simulation studies, we demonstrate LaRT's advantages over IRT in terms of superior estimation accuracy and shorter confidence intervals for latent trait estimation. To evaluate LaRT in real data, we collect responses from diverse LLMs on popular benchmark datasets. We find that LaRT yields different LLM rankings than IRT and outperforms IRT across multiple key evaluation metrics including predictive power, item efficiency, ranking validity, and LLM evaluation efficiency. Code and data are available at https://github.com/Toby-X/Latency-Response-Theory-Model

大模型评估推理能力项目反应理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。