arXiv:2606.05405cs.AIcs.CL2026-06被引 9

提出真实经济任务长周期评估基准,检验AI Agent落地能力。

Agents' Last Exam

论文配图:Agents' Last Exam
图 1 · 摘自论文原文
  • 构建覆盖1000+真实任务的长期经济任务评估框架
  • 主流模型在最难任务上平均通过率低于1%
  • 面向产业界合作设计,推动评估与实际价值对齐

近期人工智能系统在众多基准测试中取得优异表现,但这些成果并未转化为多数专业领域的经济价值部署。本文认为这一差距主要源于评估问题:现有基准缺乏对真实、高经济价值工作流的持续性能衡量。为此,本文提出「Agents' Last Exam」(ALE)基准,用于评估AI代理在长周期、高经济价值、真实世界任务中的表现,并具备可验证结果。该基准由250多位行业专家协作开发,涵盖基于美国职业分类体系O*NET/SOC 2018定义的非物理类行业,按任务分类组织为13个行业集群、55个子领域,覆盖1000多个任务。当前结果显示,最困难层级仍远未饱和:在主流基线和核心配置下,平均完整通过率不足1%。ALE被设计为持续演进的活体基准,随新工作流和行业不断扩充。更广泛而言,ALE不仅是一个排行榜,更是弥合基准表现与宏观经济影响之间鸿沟的工具。

原文摘要 · Abstract (English)

Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue that this gap is largely an evaluation problem: widely used benchmarks lack sustained performance measurement on real and economically valuable workflows. This paper introduces Agents' Last Exam (ALE), a benchmark designed to evaluate AI agents on long horizon, economically valuable, real world tasks with verifiable outcomes. Developed in collaboration with 250+ industry experts, ALE covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy). It is organized around a task taxonomy with 55 sub fields grouped into 13 industry clusters covering 1K+ tasks. Current results show that the hardest tier remains far from saturated: across mainstream harness and backbone configurations, the average full pass rate is below 1%. ALE is designed as a living benchmark: its task pool grows continuously as new workflows and industries are onboarded. More broadly, ALE is intended not merely as another leaderboard, but as an instrument for closing the gap between benchmark success and GDP relevant impact.

AI评估长周期任务经济价值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。