用少量测试题高效预测大模型在各类任务上的表现。
Cost-Efficient Estimation of General Abilities Across Benchmarks
- 结合改进的项目反应理论与自适应选题,仅需16题即可预测。
- 在112个未见任务上平均误差低于7%,仅需22000词元成本。
- 适合需要低成本评估大模型能力的研究者和工程师。
为衡量大语言模型(LLM)质量,已开发数千个多样化基准。然而已有研究显示,模型性能常可由少数潜在能力因素解释,这暗示了更高效、更系统的评估可能。我们提出以预测效度为核心标准:评估框架的优劣应基于其在有限预算下预测模型在未见任务上表现的能力。为此,我们构建了“广义项目级数据集”(WILD),包含65个模型在109,564个独特题目上对163个任务(来自27个数据集)的评估结果。该数据集首次支持在不同预算约束下,系统分析多种方法对大规模、多样化未见任务的预测能力。我们发现,将改进的多维项目反应理论(IRT)模型与基于最优实验设计的自适应选题结合,可在仅观察16个题目后,实现对112个保留基准任务的预测,平均绝对误差(MAE)低于7%。进一步引入成本感知折扣因子,使达到7% MAE所需的总词元数从141,000降至22,000,评估成本降低85%。
原文摘要 · Abstract (English)
Thousands of diverse benchmarks have been developed to measure the quality of large language models (LLMs). Yet prior work has demonstrated that LLM performance is often sufficiently explained by a small set of latent factors, or abilities. This suggests the potential for more efficient and principled benchmarking, but it remains difficult to compare the quality of different methods. Motivated by predictive validity, we argue that the quality of a benchmarking framework should be grounded in how efficiently it enables the prediction of model performance on unseen tasks. To analyze this objective, we collect the "Wide-scale Item Level Dataset" (WILD), a dataset of item-model response pairs, comprising evaluations of 65 models on 109,564 unique items spanning 163 tasks drawn from 27 datasets. This dataset enables the first analysis of how different techniques can predict a model's performance on a large, diverse collection of unseen tasks under different budget constraints. We demonstrate that combining a modified multidimensional item response theory (IRT) model with adaptive item selection driven by optimal experimental design can predict performance on 112 held-out benchmark tasks with a mean absolute error (MAE) of less than 7%, and can do so after observing only 16 items. We further demonstrate that incorporating cost-aware discount factors into our selection criteria can reduce the total tokens needed to reach 7% MAE from 141,000 tokens to only 22,000, an 85% reduction in evaluation cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。