arXiv:2608.30044cs.AIcs.CL2026-08

用语义密度重加权,让模型评估更公平、更抗重复基准干扰。

Balance of Benchmarks: Semantic Density Reweighting for Benchmark Multiplicity and Task-Conditioned Evaluation

  • 根据基准描述的语义密度分配反向权重,避免热门任务被重复计分。
  • 在586个模型14个基准上,任务特异性预测相关性达0.462,显著优于均权的0.049。
  • 可应对基准重复问题,添加4倍基准后排名相关性仍保持0.995,适合严谨评估者。

语言模型通常通过等权平均多个基准得分进行比较。然而这些基准列表随论文发布自然增长,缺乏明确测量设计,导致密集覆盖区域被反复计分,形成隐含的能力权重。本文提出平衡基准(BoB),将基准描述嵌入语义空间,并为每个基准赋予逆密度权重。相近基准在特定密度尺度下共享聚合影响。在将异构评分统一到共同潜在尺度后,利用相同几何结构构建残差场,实现基于任务查询的模型排名条件化。两个组件各司其职:在包含586个模型和14个基准的快照中,BoB能预测哪些模型在预留任务上表现异常突出,其特征相关性达0.462,远超等权下的0.049;同时有效抑制密集重复基准的影响——逐次增加每项基准四倍后,排名的肯德尔τ相关性维持在0.995,而等权下仅为0.936。残差场提供任务条件预测,逆密度加权保障对基准多重性的鲁棒性。两者共同将基准列表构成从评估套件的偶然属性,转变为可显式控制的测量设计部分,为任务感知与多重性鲁棒的模型评估提供原则性基础。

原文摘要 · Abstract (English)

Language models are commonly compared by averaging scores across a benchmark list with equal weight. Such lists grow through publication outside an explicit measurement design, so equal weighting turns the density of published benchmarks into an implicit capability weight: densely benchmarked regions count repeatedly. We introduce Balance of Benchmarks (BoB), which embeds benchmark descriptions and assigns each benchmark an inverse-density semantic weight. Nearby entries share aggregate influence at a disclosed density scale. After equating heterogeneous scores onto a common latent scale, a residual field uses the same geometry to condition model rankings on a task query. The two components serve distinct empirical roles. On a snapshot of 586 models and 14 benchmarks, BoB predicts which models are unusually strong on a held-out task beyond their general ability, reaching a profile correlation of 0.462 compared with 0.049 under equal weighting. It also limits the influence of densely repeated benchmarks on the aggregate. After adding four copies of each benchmark in turn, the resulting rankings retain a Kendall tau of 0.995, compared with 0.936 under equal weighting. The residual field therefore provides task-conditioned prediction, and inverse-density weighting provides robustness to benchmark multiplicity. Together, they turn benchmark-list composition from an incidental property of evaluation suites into an explicit, controllable part of measurement design, providing a principled foundation for task-aware and multiplicity-robust model evaluation.

模型评估基准设计语义加权鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。