arXiv:2605.25272cs.AIcs.CY2026-05

揭示大模型评测体系的隐性结构,找出排名背后的测量误差来源。

AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

  • 用因子分析和可推广性理论拆解评测分数的变异来源。
  • 发现当前评分系统存在局部依赖,且模型架构解释力弱于贡献者信息。
  • 提出更可靠的潜在能力评估方法,适合改进评测设计与信任排名。

尽管综合排行榜驱动人工智能发展,但其评分包含显著的测量噪声,其来源与大小尚不明确,导致难以判断排名反映的是真实能力差异还是评测偏差。本文提出一种测量人工智能评测生态隐性结构的框架。基于来自 Open LLM Leaderboard 的4000多个模型,运用验证性因子分析(CFA)与可推广性理论,分解排名方差来源,发现:(1) 当前报告实践中假设的基准间关系结构低估了其实际强度;(2) 排行榜项目之间存在局部依赖,削弱了基准作为测量工具的有效性;(3) 贡献者元数据在解释排名相关方差方面(约9%)优于架构或部署类别;(4) 显式分数的“缩放律”斜率可靠性低(R_β=0.53),而潜在通用因子斜率在生态系统控制下高度稳定(R_g=0.97)。研究提供了对评测动态的独特洞察,例如哪些基准随LLM规模变化,哪些反而受微调策略反向影响。我们提供可操作的诊断工具,以判断排行榜可信度并优化评测设计。

原文摘要 · Abstract (English)

While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, making it unclear when rankings reflect genuine capability differences versus evaluation artifacts. We introduce a framework for measuring the latent landscape in AI benchmark ecosystems. Applying Confirmatory Factor Analysis (CFA) and Generalizability Theory to 4,000+ models from the Open LLM Leaderboard, we decompose sources of ranking variance and establish: (1) structures assumed in current reporting practice underestimate the strength of relationships between benchmarks; (2) evidence of local dependence among leaderboard items, undermining uses of benchmarks as measurement instruments under current scoring systems; (3) contributor metadata explains more rank-relevant variance ($\approx9\%$) than architecture or deployment categories in this context; (4) a manifest-score "scaling law" slope has low reliability ($R_β=0.53$); by contrast, the latent general-factor size slope is highly stable across ecosystem controls ($R_g=0.97$). We are able to provide unique insights into benchmark dynamics, such as which benchmarks are a function of LLM size and which can be oppositely impacted by post-training practices. We provide actionable diagnostics to determine how benchmark rankings can be trusted and how benchmark design can be improved.

评测体系大模型潜变量分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。