arXiv:2607.17409stat.MLcs.LG2026-07被引 1

用历史数据高效评估新大模型性能,快速缩小置信区间。

Efficient Sequential Evaluation of Large Language Models

  • 基于逆信息投影和投注测试构建置信序列
  • 提出增长导向查询规则,加速置信区间收敛
  • 发现预测偏差与查询分布尖锐性拖慢收敛,建议混合策略

我们研究在固定问题集上,利用先前大语言模型的历史表现数据,顺序评估新大语言模型的性能。目标是构建该模型在问题集上的置信序列(CS),并设计能最快缩小CS宽度的主动查询规则。针对CS构建,我们反向使用一类检验超鞅,重点考察两种方法:基于逆信息投影(RIPr)和基于投注测试的方法。首先在理想设定下分析,证明了RIPr方法的最优性。随后提出一种增长导向查询规则,旨在最大化当前CS端点的一步期望对数增量。实际中,我们在历史数据学习到的问题级正确性预测基础上构建超鞅和查询规则。分析表明,两个关键因素会减缓置信序列收缩:累积预测偏差和查询分布的尖锐性。为此,我们提出多种混合查询规则,结合增长导向、预测优化与均匀探索,以缓解这些影响。在多个合成测试数据集上对比不同查询规则的表现,结果发现最简单的均匀采样有时反而优于更自适应的规则。

原文摘要 · Abstract (English)

We study the problem of sequentially evaluating a new large language model (LLM) on a fixed question set using historical performance data from prior LLMs. Our goal is to construct a confidence sequence (CS) for the model's capability on this question set and to design active querying rules that shrink the CS width as quickly as possible. For CS construction, we invert a family of test supermartingales and focus on two representative approaches: a reverse information projection (RIPr)-based approach and a testing-by-betting-based approach. We first study these approaches under an oracle setting, and demonstrate the oracle optimality of the RIPr-based construction. We then propose a growth-oriented querying rule that aims to maximize the worst-case one-step expected log-increment over the endpoints of the current CS. In practice, we build these test supermartingales and the querying rule on predictions of question-level correctness learned from historical data. We then analyze the shrinkage behavior of the resulting CSs and identify two key factors that slow the shrinkage rate of CSs: accumulated prediction mismatch and the spikiness of the querying distribution. Finally, motivated by this analysis, we propose several mixture querying rules that combine growth-oriented querying, prediction refinement, and uniform exploration, trying to mitigate the effects that slow the shrinkage rate. We provide experiments comparing different querying rules for the RIPr-based and testing-by-betting-based CSs across several synthetic testing datasets. Interestingly, we observe that the simplest querying rule, uniform sampling, can sometimes outperform more adaptive querying rules for both methods.

大模型评估置信序列主动学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。