arXiv:2603.29357cs.AI2026-03被引 2

用有效维度检测评测集冗余,发现多数评分其实只反映少量独立信息。

BenchScope: How Many Independent Signals Does Your Benchmark Provide?

  • 提出有效维度(ED)衡量评测分数的独立信号数量,快速诊断测量广度。
  • 实测显示开放大模型榜单仅相当于2个有效维度,多个评测高度相似。
  • 适合评测设计者使用,可识别冗余项目并优化评测体系结构。

AI评估套件常报告多个分数却未验证其是否提供独立信息。本文引入有效维度(ED),即中心化评测分数谱的参与比率,作为测量广度的快速、群体条件上界诊断工具。在8个领域、22个评测、超过8,400次模型评估中应用ED,发现显著冗余:六分制的Open LLM Leaderboard仅相当于约两个有效测量轴(ED = 1.7),BBH与MMLU-Pro高度可互换(ρ = 0.96,跨七类子群体稳定),当前评测的测量广度差异超20倍。我们证明相对ED排名在匹配维度控制下稳定,并展示ED可识别冗余组件、监测性能条件压缩、指导评测维护。由于二元谱会高估潜在维度,将ED视为筛选统计量而非真实因子数,辅以零假设、可靠性与饱和性分析。我们提供22个评测的参考图谱及四步诊断工作流,只需输入分数矩阵与少量代码即可运行。

原文摘要 · Abstract (English)

AI evaluation suites often report many scores without checking whether those scores carry independent information. We introduce Effective Dimensionality (ED), the participation ratio of a centered benchmark-score spectrum, as a fast, population-conditional upper-bound diagnostic of measurement breadth. Applied at per-instance granularity to 22 benchmarks across 8 domains and more than 8,400 model evaluations, ED reveals substantial redundancy: the six-score Open LLM Leaderboard behaves like roughly two effective measurement axes (ED = 1.7), BBH and MMLU-Pro are near-interchangeable (rho = 0.96, stable across seven subpopulations), and measurement breadth varies more than 20x across current benchmarks. We show that relative ED rankings are stable under matched-dimension controls and that ED can flag redundant suite components, monitor performance-conditional compression, and guide benchmark maintenance. Because binary spectra overestimate absolute latent dimensionality, we interpret ED as a screening statistic rather than a literal factor count and complement it with null, reliability, and saturation analyses. We provide a 22-benchmark reference atlas and a four-step diagnostic workflow that benchmark maintainers can run with a score matrix and a few lines of code.

评测设计有效维度冗余检测基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。