解决大模型评估中因自适应测试导致的高估问题,提升结果可靠性。
Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking

- 提出SIREN协议,分离选择与评估,冻结候选集防止过拟合。
- 使用项级自助法量化不确定性,在有限预算下保证推断有效。
- 适用于需要可靠性能比较的模型调优与部署场景。
自适应提示与程序搜索使大模型评估对数据选择敏感。一旦基准项在调优过程中被重复使用,观察到的最优得分无法准确反映完整‘调优后部署’流程的真实表现。本文研究在明确调优预算下的过程级目标推断。提出SIREN——一种选择感知的重复划分报告协议:冻结搜索后短名单,分离逐次选择与保留集评估,并采用项级高斯乘子自助法进行不确定性量化。在固定短名单、选择稳定的条件下,该估计器具有首阶项级表示形式,自助法可在有限预算网格上实现有效的联合推断。支持过程性能曲线的置信区间及预设等预算与跨预算比较。控制模拟与MMLU-Pro调优实验表明,基于胜者的结果报告可能过于乐观并改变部署结论,而SIREN始终贴近有限样本报告目标。
原文摘要 · Abstract (English)
Adaptive prompt and program search makes LLM evaluation selection-sensitive. Once benchmark items are reused inside tuning, the observed winner's score need not estimate the fresh-data performance of the full tune-then-deploy procedure. We study inference for this procedure-level target under explicit tuning budgets. We propose SIREN, a selection-aware repeated-split reporting protocol that freezes the post-search shortlist, separates splitwise selection from held-out evaluation, and uses an item-level Gaussian multiplier bootstrap for uncertainty quantification. In a fixed-shortlist regime with smooth stabilized selection, the estimator admits a first-order item-level representation, and the bootstrap yields valid simultaneous inference on a finite budget grid. This supports confidence intervals for procedure-performance curves and pre-specified equal-budget and cross-budget comparisons. Controlled simulations and MMLU-Pro tuning experiments show that winner-based reporting can be optimistic and can change deployment conclusions, while SIREN remains close to the finite-sample reporting target.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。