arXiv:2608.29420cs.LGcs.CY2026-09

经济类评测虽有额外信息,但主要反映模型发布时间差异。

One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation

论文配图:One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation
图 1 · 摘自论文原文
  • 用潜在变量模型分析421个模型在12项评测中的表现
  • 单一因素解释74.5%变异,且与发布时间强相关(R²=0.505)
  • 经济评测不构成独立能力维度,但能补充时间趋势外的信息

前沿模型排行榜现以经济基准测试为依据,评估模型在软件工程、银行流程等专业任务中的表现,这些排名影响组织采购、监管审查及对工作变革的预期。然而,此类基准是否衡量一种独立于通用测试能力的特性,还是仅反映模型整体进步的单一维度,尚未被研究。我们基于421个模型配置在12项基准(4项经济类)的快照数据,将基准视为项目、模型视为应答者,采用预设阈值的潜在变量模型检验四种假设。结果表明,单一因子解释了74.5%的共变,且与模型发布日期高度相关(R² = 0.505)。移除时间趋势后,该比例下降14.9个百分点,若按基础模型去重则下降24.1个百分点。根据预先设定的维度规则,经济基准未形成独立因子;但通过逐折留一法重新估计因子,多因子模型在预测保留的经济评分上优于单一般指数(合并Δ-MSE = 0.037,95%置信区间[0.019, 0.055])。因此,经济基准虽具增量预测价值,但不足以构成独立潜能力,其核心仍受时间趋势驱动。排行榜仍可反映总体进展,但短期内模型差距多源于日历时间,需校正时间偏差后才能判断真实能力差异。

原文摘要 · Abstract (English)

Frontier-model leaderboards now rank systems based on economic benchmarks, tests of how well models carry out professional tasks from software engineering to banking workflows, and those rankings inform what organisations buy, what regulators scrutinise, and expectations of how work will change. Whether such benchmarks measure a capability distinct from general test-taking, or re-express the one axis along which every benchmark rises as models improve, is a question of construct validity that has not yet been studied. We test it on a hash-pinned leaderboard snapshot of 421 model configurations across twelve benchmarks, four of them economic, treating benchmarks as items and models as respondents in a latent-variable model with four hypotheses and their thresholds fixed before analysis. A single factor explains 74.5% of common variance and tracks model release date (R^2 = 0.505), so the leading axis of capability is substantially a time trend; where prior work controls for scale, compute adds little once date is removed. Removing the date trend lowers that share by 14.9 points, and by 24.1 with one row per base model. Under the dimensionality rule fixed in advance the economic benchmarks form no distinct factor, yet a leave-one-benchmark-out test with factors re-estimated inside every fold shows that a multi-factor representation predicts held-out economic scores better than a single general index (pooled Delta-MSE 0.037, 95% bootstrap interval [0.019, 0.055]). Economic benchmarks therefore add incremental predictive information to a largely date-driven general factor, and the evidence does not support treating them as a distinct latent capability. Leaderboards remain a sound guide to overall progress, but most of the gap between models released months apart is calendar, so a small gap between contemporaneous models should be date-adjusted before being read as a capability difference.

模型评估经济基准时间趋势潜在变量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。