arXiv:2606.08679stat.MLcs.CL2026-06被引 1

给模型排名加不确定性区间,让排名更可靠。

Rank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation

  • 分层框架结合配对比较与置信推断,量化每项任务的排名置信区间。
  • 在TabArena和MMLU上验证,排名区间具有统计有效性且信息丰富。
  • 适合关注模型性能不确定性的研究者和评估者使用。

预训练模型常通过多任务排行榜评估其在不同场景下的适用性。然而,现有方法在将各任务表现聚合为排行榜排名时,未考虑任务层面的不确定性与变异性。尽管已有工作提出基于区间的模型排名,但如何从单个任务的不确定性合理推导出排行榜层级的排名区间仍缺乏系统方法,且模型在不同任务上的表现差异常被掩盖。本文提出一种分层框架,构建具有统计保障的模型排名区间:在任务层级,基于配对比较生成排名置信区间;在排行榜层级,采用约等于方法生成排名预测区间。该方法可对已观测任务及潜在新任务实现可靠的模型排名不确定性量化。在模拟数据以及TabArena和PromptEval(MMLU)基准上的实验表明,本方法产生的区间具有统计有效性且信息量充足,支持更可靠的、带有不确定性的排行榜模型评估。

原文摘要 · Abstract (English)

Pretrained models are often evaluated on multi-task leaderboards to measure their applicability in diverse contexts. However, current methods for aggregating performance across tasks into leaderboard-level rankings do not address the uncertainty and variability at the task level. While recent works have proposed interval-based model rankings, the principled aggregation of uncertainty from individual tasks to leaderboard-level rankings remains unaddressed, and variation in models' performance across tasks is frequently obscured. In this work, we introduce a hierarchical framework that constructs model rank intervals with statistical guarantees at both levels: task-level rank confidence intervals from pairwise comparisons, and leaderboard-level rank prediction intervals using a conformal approach. This enables reliable quantification of model rank for each observed task and for new potential tasks. Experiments on simulated data and the TabArena and PromptEval (MMLU) benchmarks show that our method yields statistically valid and informative intervals, enabling reliable, uncertainty-aware model ranking on leaderboards.

模型评估不确定性排行榜置信区间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。