构建远程工作自动化评估基准,实测AI仅能自动化2.5%任务。
Remote Labor Index: Measuring AI Automation of Remote Work
- 设计跨行业真实项目评测框架,衡量AI端到端执行能力。
- 顶尖AI agent仅实现2.5%的自动化率,远低于预期。
- 为劳动力自动化提供可量化的实证依据,适合政策与企业参考。
尽管AI在知识与推理类研究基准上进展迅速,但其经济价值与自动化潜力仍不明确。为此,我们提出远程工作自动化指数(Remote Labor Index, RLI),一个涵盖多行业的现实世界任务基准,用于评估AI代理在实际场景中的端到端表现。实验显示,当前最先进AI代理在RLI上的自动化率仅为2.5%,接近最低水平。该结果为人工智能自动化讨论提供了实证基础,有助于各方跟踪技术影响并主动应对自动化带来的劳动力变革。
原文摘要 · Abstract (English)
AIs have made rapid progress on research-oriented benchmarks of knowledge and reasoning, but it remains unclear how these gains translate into economic value and automation. To measure this, we introduce the Remote Labor Index (RLI), a broadly multi-sector benchmark comprising real-world, economically valuable projects designed to evaluate end-to-end agent performance in practical settings. AI agents perform near the floor on RLI, with the highest-performing agent achieving an automation rate of 2.5%. These results help ground discussions of AI automation in empirical evidence, setting a common basis for tracking AI impacts and enabling stakeholders to proactively navigate AI-driven labor automation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。