量化大模型评测中排名的不确定性,揭示成绩波动原因
Quantifying Ranking Uncertainty in LLM Benchmarks
- 用配对假设检验聚合方法分析评测不确定性来源
- 发现MMLU各科目排名差异显著,跨科目比较需谨慎
- 为模型选型和性能评估提供可信区间参考
预训练模型通常通过多任务排行榜评估其在各类任务中的表现。近期引入了排名置信区间,通过聚合配对假设检验来量化这些排名的不确定性。本文分析了知识评估基准MMLU中的不确定性来源,并展示了如何修改假设检验以考虑这些影响。结果表明,大语言模型在MMLU不同科目间的排名波动显著,因此在比较模型性能或识别最优模型时必须考虑这一因素。
原文摘要 · Abstract (English)
Pretrained models are typically ranked on multi-task leaderboards to assess their effectiveness across diverse tasks. Rank confidence intervals were recently introduced as a method to quantify the uncertainty in these rankings by aggregating pairwise hypothesis tests. In this work, we analyze the sources of uncertainty in the knowledge evaluation benchmark MMLU and show how hypothesis tests can be modified to account for their effects. We demonstrate that ranking variability across MMLU subjects is substantial and should be considered when comparing LLMs or identifying the top-performing models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。