arXiv:2502.15850cs.CLcs.AI2025-02被引 5

预测大模型智能体能力,为社会应对提供参考

Forecasting Frontier Language Model Agent Capabilities

  • 分两步预测:先估评分再推基准表现
  • 2026年普通模型在代码任务上达54%成功率
  • 适合关注AI发展节奏的政策与技术决策者

随着语言模型日益作为自主智能体运行,准确预测其能力对社会准备至关重要。我们评估了六种预测方法,分为‘一步’(直接由计算资源或发布日期预测基准得分)和‘两步’(先预测跨基准性能主成分PC-1和人工评估的竞争力Elo评分)。通过在OpenLLM 2排行榜上的38个模型数据进行回测,验证了‘发布日期→Elo→基准’的两步法。据此预测,到2026年初,未专门优化的模型在SWE-Bench Verified(软件开发)上成功率可达54%,而顶尖模型将达87%。该方法未考虑近期推理-计算扩展进展,可能偏保守。

原文摘要 · Abstract (English)

As Language Models (LMs) increasingly operate as autonomous agents, accurately forecasting their capabilities becomes crucial for societal preparedness. We evaluate six forecasting methods that predict downstream capabilities of LM agents. We use "one-step" approaches that predict benchmark scores from input metrics like compute or model release date directly or "two-step" approaches that first predict an intermediate metric like the principal component of cross-benchmark performance (PC-1) and human-evaluated competitive Elo ratings. We evaluate our forecasting methods by backtesting them on a dataset of 38 LMs from the OpenLLM 2 leaderboard. We then use the validated two-step approach (Release Date$\to$Elo$\to$Benchmark) to predict LM agent performance for frontier models on three benchmarks: SWE-Bench Verified (software development), Cybench (cybersecurity assessment), and RE-Bench (ML research engineering). Our forecast predicts that by the beginning of 2026, non-specialized LM agents with low capability elicitation will reach a success rate of 54% on SWE-Bench Verified, while state-of-the-art LM agents will reach an 87% success rate. Our approach does not account for recent advances in inference-compute scaling and might thus be too conservative.

大模型预测智能体能力前瞻评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。