揭示大模型评测的隐藏盲区,发现评测集覆盖不足是排名失真主因。
The Evaluation Blind Spot: A Stereological Theory of Benchmark Coverage for Large Language Models
- 用立体几何理论分析评测集覆盖范围,量化盲区大小。
- 实测三大榜单盲区比排名差距大100倍以上,远超随机噪声。
- 提出稳定评测子集构建方法,4个基准可覆盖90%能力维度。
本文建立大语言模型评测覆盖的立体几何理论。对于有效维度为d_eff的评测集,两个一致得分的凸能力轮廓间可见豪斯多夫距离受epsilon + C R m^(-1/(d_eff-1))限制,并有匹配的利普希茨下界。实证显示,Open LLM v2、扩展12项评测集和LiveBench三个排行榜在竞争前沿的d_eff介于[2.86, 4.80];结构性盲区超出观测亚军分差两个数量级,且压倒统计噪声52–127倍。在卡方投影模型下,各向同性先验为乐观情况;六种隐含能力先验与四种环境维度下,前两名模型的半分裂互换率稳定在[0.38, 0.49],500次随机可见/保留分割试验中92%出现第一名互换,平均5个中有2.83个顶级模型变更。采用子模贪心算法(具Nemhauser (1 - 1/e)保证)找到稳定核心4个基准;12个中7个即可实现90%覆盖,且训练子集在时间季度间转移时保持93–97%性能保留。12个内部评测与27个Chatbot Arena类别的反事实验证表明,特征结构能准确预测哪些评测不可替代(ρ = -0.69, p = 0.013),哪些外部评测带来新信息(ρ = +0.38)。第二项独立理论贡献为解决Gardner问题1.5(1995),在一般维数下通过S^(D-1)上的最优恢复理论确立最小最大率Theta(R/(kappa m^(2/(D-1))))。
原文摘要 · Abstract (English)
We give a stereological theory of LLM benchmark coverage. For any suite with effective dimensionality d_eff, the visible Hausdorff distance between two convex capability profiles consistent with the same scores is bounded by epsilon + C R m^(-1/(d_eff-1)), with matching Lipschitz lower bound. Empirically, three independent leaderboards (Open LLM v2, an extended 12-benchmark suite, LiveBench) all have d_eff in [2.86, 4.80] on their competitive frontier; the structural blind spot exceeds the observed runner-up score gap by two orders of magnitude and dominates statistical noise by 52-127x. Under a chi-squared projection model, the isotropic prior is the optimistic case; across six hidden-capability priors and four ambient dimensions the simulated half-split swap rate of the top two models stays in [0.38, 0.49], and a 500-trial random visible/held-out split shows that 92% of trials swap the top-1 ranking with on average 2.83 of 5 top-5 models changing. A submodular greedy algorithm with the Nemhauser (1 - 1/e) guarantee finds a stable core of 4 benchmarks; 7 of 12 suffice for 90% coverage, and the trained subset transfers across temporal quarters with 93-97% retention. A counterfactual validation across 12 internal benchmarks and 27 Chatbot Arena categories confirms that the eigenstructure predicts which evaluations are irreplaceable (rho = -0.69, p = 0.013 for removal disruption) and which external evaluations bring new information (rho = +0.38). As a second, independent theoretical contribution, we resolve Gardner's Problem 1.5 (1995) for C^2 support functions, establishing the minimax rate Theta(R/(kappa m^(2/(D-1)))) in general dimension via optimal recovery theory on S^(D-1).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。