arXiv:2608.03437cs.CLcs.LG2026-08

用自适应采样减少人工评估成本,更快找出表现最好的模型

Dynamically Allocating Evaluation Effort for Model Ranking

  • 根据已有评分动态聚焦评估顶尖模型,避免盲目全量评测
  • 相比传统方法,相同预算下能更清晰区分顶尖模型性能
  • 适合大规模模型竞赛或资源有限的评测场景

尽管人工评估是许多NLP任务的黄金标准,但其成本高昂且难以扩展。当前典型评估协议会耗尽全部资源对所有模型在完整基准上进行评价,虽安全却效率低下。本文将多模型人工评估形式化为具有相关臂的多臂老虎机中的最优臂识别问题,其中每次抽臂对应一次人工评估。通过基于已有样本的中间排名结果自适应采样,可将标注预算集中于最具竞争力的模型。我们证明了所提算法的最优性,并表明其显著提升了对顶尖模型间的区分能力,使评估更快速、更经济,也更符合大规模竞赛的评估目标。

原文摘要 · Abstract (English)

While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing models, typical evaluation protocols waste effort by exhaustively evaluating all models on the entire benchmark, a safe but inefficient approach. In this work, we formalize multi-model human evaluation as a best-arm identification problem in a multi-armed bandit setup with correlated arms, where pulling an arm corresponds to human-evaluating a model. By sampling adaptively based on the intermediate model rankings obtained on the samples so far, we can focus the annotation budget on the most competitive models. We prove the optimality of the proposed algorithms and show that it improves discrimination between top-performing models. This makes evaluations faster, cheaper and more aligned with large-scale competition evaluation goals.

模型评估多臂老虎机人工标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。