用动态评估取代固定样本,高效又可靠地判断模型好坏。
Stop Guessing When to Stop Testing: Efficient Model Evaluation with Just Enough Data

- 引入序贯检验框架,根据实际需求动态决定测试数据量。
- 在开放视觉语言模型榜单上降低80%计算成本,误差区间仅放宽2.5点。
- 适合需要频繁评估模型的开发者,尤其关注效率与精度平衡者。
固定大小的评估基准固有的僵化性使其在模型评估中效率低下。不同的评估目标,如模型排序、选择及开发过程中的持续测试,需要不同程度的统计功效。固定样本量与这些多样化需求之间的不匹配,导致计算成本过高或可靠性下降,严重影响模型评估质量。为此,我们主张在该领域采用序贯测试。本文提出一种自适应评估框架,为评估中的效率与可靠性权衡提供系统方法。该框架结合成熟的序贯检验统计范式,设计了针对常见评估需求的停止准则,例如检测收益递减和最小可检测效应。我们在 Open VLM Leaderboard 上验证了其能力,例如相比固定样本评估,计算成本降低80%(允许2.5点置信区间宽度),同时保持统计显著性。
原文摘要 · Abstract (English)
The inherent rigidity of fixed-size benchmarks makes them an inefficient tool for model evaluation. Diverse evaluation objectives, including model ranking, model selection and testing throughout development, demand varying levels of statistical power. The mismatch between fixed sample sizes and these diverse needs results in either excessive computational cost or compromised reliability - a critical concern for model evaluation. To overcome these limitations, we call for adoption of sequential testing in our field. We provide an adaptive evaluation framework, that provides a principled way to navigate the trade-off between efficiency and reliability in model evaluation. Our framework combines the established statistical paradigm of sequential testing with stopping criteria tailored to common evaluation needs such as diminishing returns detection, and minimum detectable effect size. We demonstrate its ability to adaptively manage the efficiency-reliability trade-off on the Open VLM Leaderboard, including, for example, a 80% reduction in computational cost compared to fixed-size evaluation (with a 2.5-point CI width allowance) while maintaining statistical significance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。