用固定题库让大模型评测随时间扩展,结果仍可比。
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration

- 用锚题固定参数,新数据集加入时只需测锚题。
- 仅用100道锚题即可预测完整评测结果,误差2-3个百分点。
- 适合长期跟踪模型性能的研究者与评测平台。
语言模型和评测集的快速迭代使得对每个模型在每个数据集上进行评估变得成本高昂。实际中,模型常在不同样本上测试,导致跨研究结果难以比较。为此,我们提出基于多维项目反应理论(IRT)的框架,利用锚题将新评测集校准到现有评估体系,同时保持先前校准的题目参数固定。该方法支持数据集随时间逐步引入的现实评估场景,模型仅在评估时可用的数据集上测试,通过每数据集固定的锚题集合实现不同时期结果的直接比较。在超过400个模型的大规模实验中,该框架仅使用每数据集100个锚题,即可在2-3个百分点内预测完整评估表现,排名保留的斯皮尔曼相关系数ρ≥0.9。这表明评测套件可随时间扩展,且分数仍具可比性——新增数据集只需运行已有模型在该数据集的锚题上即可。
原文摘要 · Abstract (English)
The rapid release of both language models and benchmarks makes it increasingly costly to evaluate every model on every dataset. In practice, models are often evaluated on different samples, making scores difficult to compare across studies. To address this, we propose a framework based on multidimensional Item Response Theory (IRT) that uses anchor items to calibrate new benchmarks to the evaluation suite while holding previously calibrated item parameters fixed. Our approach supports a realistic evaluation setting in which datasets are introduced over time and models are evaluated only on the datasets available at the time of evaluation, while a fixed anchor set for each dataset is used so that results from different evaluation periods can be compared directly. In large-scale experiments on more than 400 models, our framework predicts full-evaluation performance within 2-3 percentage points using only 100 anchor questions per dataset, with Spearman $ρ\geq 0.9$ for ranking preservation preservation. This shows that benchmark suites can grow over time while preserving score comparability, since adding a new dataset requires running existing models only on that dataset's anchors. Code and data are available at: https://eliyahabba.github.io/growing-pains/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。