现有前沿AI预测缺乏可靠数据支撑,需重建可审计的测量体系。
Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence

- 构建截至2026年8月12日的事件驱动型测量记录,覆盖62个系统与144个评估事件
- 仅7个系统同时具备训练算力与50%任务达成时间数据,多数闭源系统缺失关键信息
- 发现基准迭代存在非线性跃迁,现有方法对趋势判断统计效能不足80%
前沿人工智能的定量预测常将过时目标与基准得分、训练算力、发布时间或专家意见趋势关联。本文审计公开测量记录是否支持此类关联,再拟合新趋势。构建截至2026年8月12日的冻结事件中心记录,包含62个选定系统、12个版本化基准、7项能力/影响标准、144个分级事件、27个来源记录及408种类型关系。该记录为审计样本,非普查。仅有7个系统同时观测到训练算力与METR 50%任务时间阈值;27个闭源系统中19个缺失算力数据,2026年所有闭源发布均无此数据;35个开源权重系统均无METR时间阈值观测。基准更迭引发二次断裂:从METR 1.0到1.1的七系统连接在对数尺度上斜率为1.206(95%置信区间1.021至1.390),而六系统MMLU到MMLU-Pro比较在逻辑斯蒂和概率链接下呈现跳跃式变化,但在线性或对数链接下不显著。现有桥梁对近25%斜率偏移的检测功效约80%。来源高度集中:71个实质性量化事件中52个(73.2%)来自单一测量计划,76.1%为实验室发布。对56项方法论与实证文献的审查识别出16个互补测量方向,涵盖资源、推理预算、可靠性、智能体工作、安全、人类偏好、实际成果及预测回测等,但无一提供可替代标量。结论并非前沿AI预测不可能,而是可信的日期预测必须明示其依赖的版本化测量系统、显式连接、协议、链接与来源依赖,而非仅依赖拟合曲线或日历日期。
原文摘要 · Abstract (English)
Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time, or expert belief. This paper audits whether the public measurement record supports those connections before another trend is fitted. I construct a frozen, event-centric record through 12 August 2026 with 62 selected systems, 12 versioned benchmarks, seven capability or impact criteria, 144 graded events, 27 source records, and 408 typed relations. The record is an audit sample, not a census. Only seven systems jointly observe estimated training compute and a METR 50 percent task horizon. Training compute is absent for 19 of 27 closed systems, including every selected closed release from 2026, while none of the 35 open-weight systems has a METR horizon observation. Benchmark succession creates a second break: a seven-system link from METR Time Horizon 1.0 to 1.1 has a log-scale slope of 1.206 (95 percent CI 1.021 to 1.390), whereas a six-system MMLU to MMLU-Pro comparison appears shift-like under logit and probit links but not under linear or logarithmic links. The observed bridges have about 80 percent power only for slope departures near 25 percent. Provenance is concentrated: 52 of 71 substantive quantitative events, or 73.2 percent, come from one measurement programme, and 76.1 percent are laboratory releases. A review of 56 methodological and empirical sources identifies 16 complementary measurement directions spanning resources, inference budgets, reliability, agentic work, safety, human preference, field outcomes, and forecast backtesting. No direction supplies a replacement scalar. The result is not that frontier AI forecasting is impossible, but that a defensible dated forecast is a claim about a versioned measurement system with explicit joins, protocols, links, and source dependence, not merely a fitted curve or calendar date.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。