公开AI评估数据存疑,新方法用贝叶斯推理揭露时间假象
Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations
- 用贝叶斯框架分析评估历史,揭示同一结果可对应不同时间线
- 实测显示达到上限时间可能为23.03或75.13,差距超三倍
- 适合关注模型评估可信度的研究者与审稿人
公开的AI评估常被视为最终排名,但其背后证据是受报告规则、基准更新和缺失数据影响的选择性时间序列。LiveBench和Open LLM Leaderboard v2的重复归档构成主要纵向记录;LMArena提供偏好压力测试;GAIA和tau-bench贡献有限的代理型评测。这些档案共同构成一个贝叶斯推断问题:在固定报告约定下,仅基于终端结果的1,000个系统实例,存在两种前终端历史路径,导致达到上限0.05内的预测时间分别为23.03或75.13。合成后验对比显示,面向行动的诊断结果随观测方式而异。候选选择感知的前沿模型在合成恢复、客观档案预测、偏好迁移和不确定性校准上均失败,相应地,固定审计门限拒绝其更强主张。一种档案-裁决协议可重建评估历史,隔离验证的时间边界,并驳回未经支持的前沿声称。
原文摘要 · Abstract (English)
Public AI evaluations are often read as terminal leaderboards, yet the underlying evidence is a selective time series shaped by reporting rules, benchmark revisions, and missingness. Repeated public archives for LiveBench and Open LLM Leaderboard v2 serve as the primary longitudinal record; LMArena provides a preference stress test; and GAIA and tau-bench contribute limited agentic pilots. Together, these archives instantiate a Bayesian inference problem: under a fixed reporting convention, one constructed terminal-only example over $1{,}000$ systems is compatible with two pre-terminal histories, yielding times of $23.03$ or $75.13$ to reach within $0.05$ of the ceiling under the same terminal-tail model. In synthetic posterior comparisons, action-facing diagnostics differ across observation regimes. The candidate selection-aware frontier model fails synthetic recovery, objective-archive prediction, preference transfer, and uncertainty calibration; correspondingly, fixed audit gates reject its stronger claims. An archive-and-adjudication protocol reconstructs public evaluation histories, isolates a verified timing boundary, and falsifies unsupported frontier claims.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。