arXiv:2509.24086cs.AIcs.CL2025-09被引 16

单次评估易出错,至少需两次重复才能可靠比较大模型性能。

Do Repetitions Matter? Strengthening Reliability in LLM Evaluations

  • 通过多次独立运行评估模型,减少随机波动对排名的影响。
  • 单次运行时83%的评估片段会出现排名反转,两次运行可消除83%错误。
  • 建议采用不少于两次重复评估,提升结果可信度,适合小团队实践。

大模型排行榜通常依赖单次随机运行,但所需重复次数以确保结论可靠尚不明确。我们对八个顶尖模型在AI4Math基准上进行了每种设置三个独立运行的重新评估。采用混合效应逻辑回归、领域级边际均值、排名不稳定性分析及运行间可靠性分析,评估额外重复的价值。研究发现:单次运行的排行榜极不稳定——12个评估片段中有10个(83%)在与三次运行多数结果对比时发生至少一次配对排名反转,尽管成对显著性无符号变化且组间相关性中等。平均运行结果仅使标准误缩小约5%(从一次到三次),但排名改善显著;两次运行即可消除约83%的单次运行排名反转。我们为实践者提供成本敏感的建议:将评估视为实验,报告不确定性,并在随机解码下使用不少于两次重复。这些做法在保持可行性的同时提升鲁棒性,更贴近真实世界可靠性。

原文摘要 · Abstract (English)

LLM leaderboards often rely on single stochastic runs, but how many repetitions are required for reliable conclusions remains unclear. We re-evaluate eight state-of-the-art models on the AI4Math Benchmark with three independent runs per setting. Using mixed-effects logistic regression, domain-level marginal means, rank-instability analysis, and run-to-run reliability, we assessed the value of additional repetitions. Our findings shows that Single-run leaderboards are brittle: 10/12 slices (83\%) invert at least one pairwise rank relative to the three-run majority, despite a zero sign-flip rate for pairwise significance and moderate overall interclass correlation. Averaging runs yields modest SE shrinkage ($\sim$5\% from one to three) but large ranking gains; two runs remove $\sim$83\% of single-run inversions. We provide cost-aware guidance for practitioners: treat evaluation as an experiment, report uncertainty, and use $\geq 2$ repetitions under stochastic decoding. These practices improve robustness while remaining feasible for small teams and help align model comparisons with real-world reliability.

大模型评估可靠性重复实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。