提出评估大模型配对比较的分辨率诊断方法,发现多数排名结果统计不显著。
Resolution Diagnostics for Paired LLM Evaluation
- 将大模型配对评估建模为假设检验,用分辨比q诊断每对结果是否可靠。
- 40组对比中11组、MMLU-Pro中6组未达统计显著性标准,真实数据下仍不显著。
- 揭示常用计算工具在近似比较中误差翻倍,适合关注评估严谨性的研究者。
在两个公开的大模型排行榜中,许多配对排名未达到常规配对检验的分辨率目标:40组Open LLM Leaderboard v1对比中有11组、9组MMLU-Pro前10名相邻对中有4组未达(α, 1-β) = (0.05, 0.8)标准。在真实学科聚类下,该数字升至6/9,并在99.9%的类别自助重采样中保持5-6/9。本文将配对大模型评估视为假设检验问题,反向使用α水平与1-β功效检验,以每对分辨比q = N/N*作为主要诊断指标。通过显式二阶常数的小效应扩展显示,广泛使用的非配对Cohen-h-plus-(1-rho)简化公式在接近比较场景下与正确样本量N*相差约两倍,这一偏差被三个现成计算器(Cohen 1988, G*Power, R pwr)在用户对单臂输出乘以(1-rho)时隐性继承。即使经过多重性校正和任意时间有效序贯检验,未解决配对模式依然存在。
原文摘要 · Abstract (English)
Across two public LLM leaderboards, many displayed pairwise rankings do not meet a conventional paired-test resolution target under the actual paired evaluation design: 11 of 40 Open LLM Leaderboard v1 pairwise comparisons and 4 of 9 MMLU-Pro top-10 adjacent-rank pairs are unresolved at (alpha, 1-beta) = (0.05, 0.8). The MMLU-Pro count rises to 6/9 under real subject-level clustering and stays at 5-6 out of 9 in 99.9% of category-bootstrap resamples. We frame paired LLM evaluation as a hypothesis-testing problem, invert level-alpha, power-(1-beta) tests, and report a per-pair resolution ratio q = N/N* as the primary diagnostic. A sharp small-effect expansion with an explicit second-order constant shows that the widely-used unpaired Cohen-h-plus-(1-rho) shortcut deviates from the correct N* by approximately a factor of two in the close-comparison regime, a deficit that three of five off-the-shelf calculators(Cohen 1988, G*Power, R pwr) silently inherit when the user post-multiplies their per-arm output by (1-rho). The unresolved-pair pattern remains under multiplicity correction and anytime-valid sequential testing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。