arXiv:2605.20520cs.AI2026-05被引 7

用真实世界任务评估前沿AI,发现潜在的广泛能力。

Open-World Evaluations for Measuring Frontier AI Capabilities

论文配图:Open-World Evaluations for Measuring Frontier AI Capabilities
图 1 · 摘自论文原文
  • 以长期、复杂的真实任务代替传统基准测试。
  • AI成功开发并上线iOS应用,仅需一次手动干预。
  • 适合关注AI实际落地能力的研究者与从业者。

基于基准的评估在追踪前沿AI进展中仍具重要性,但其往往高估或低估实际部署能力,因偏爱可精确定义、自动评分、易优化且低成本短周期的任务。本文倡导一种互补的评估方式——开放世界评估:通过小样本定性分析,考察长期、混乱、贴近现实的任务。我们综述了近期开放世界评估,指出其优缺点,并提出CRUX(协作研究更新AI预期)项目,定期开展此类评估。作为首例,我们让AI代理自主开发并发布一个简单iOS应用至苹果应用商店。该代理仅因一次可避免的手动干预失败,表明开放世界评估能提前预警即将普及的AI能力。最后,本文提出设计和报告开放世界评估的建议。

原文摘要 · Abstract (English)

Benchmark-based evaluation remains important for tracking frontier AI progress. But it can both overstate and understate deployed capability because it privileges tasks that can be precisely specified, automatically graded, easy to optimize for, and run with low budgets and short time horizons. We advocate for a complementary class of evaluations, which we term open-world evaluations: long-horizon, messy, real-world tasks assessed through small-sample qualitative analysis rather than benchmark-scale automation. In this paper we survey recent open-world evaluations, identify their strengths and limitations, and introduce CRUX (Collaborative Research for Updating AI eXpectations), a project for conducting such evaluations regularly. As a first instance, we task an AI agent with developing and publishing a simple iOS application to the Apple App Store. The agent completed the task with only a single avoidable manual intervention, suggesting that open-world evaluations can provide early warning of capabilities that may soon become widespread. We conclude with recommendations for designing and reporting open-world evals.

AI评估开放世界实测能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。