两年间工作代理性能飙升至98%完成率,安全性和效率同步提升。
WorkBench Revisited: Workplace Agents Two Years On
- 用Claude Fable 5验证新基准,任务完成率从43%跃升至98%
- 错误行为率由26%降至1.9%,能力与安全不再冲突
- 开源模型降低成本,但基础错误仍可能导致不可逆损害
2024年3月,WorkBench上表现最佳的代理GPT-4仅完成43%的任务。我们于2026年6月重新评估该基准,发现当前表现最优的代理Claude Fable 5已实现98%的任务完成率。除性能显著提升外,三个现象尤为突出:首先,非预期有害行为(如发错邮件)占比从GPT-4的26%下降至Claude Fable 5的1.9%,表明能力与安全性在WorkBench上协同提升而非此消彼长;其次,开放权重模型的兴起大幅降低了达到高性能所需的成本,而前沿模型成本则保持稳定;最后,尽管多种错误类型已被消除,但前沿模型仍会犯一些基础性错误,偶尔引发不可逆损害。我们发布了更新版基准,包含数据与代码质量改进、新模型得分及自2024年以来代理进展的分析。
原文摘要 · Abstract (English)
The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best agent to date, Claude Fable 5, now completes 98%. Beyond this considerable progress in frontier agent performance, three things stand out. First, unintended harmful actions, such as emailing the wrong person, fell from 26% of tasks for GPT-4 to 1.9% for Claude Fable 5; capability and safety go together on WorkBench rather than trade off, so the models that finish the most tasks also do the least unintended damage. Second, the rise of open-weight models has drastically lowered costs for a performance level that was only accessible to proprietary models, while frontier costs have stayed stable. Third, while several classes of error have been eliminated, frontier models still make some basic mistakes that occasionally result in irreversible harm. We release an updated version of the benchmark with data and code quality improvements, new model scores, and analysis of agent progress on WorkBench since 2024.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。