arXiv:2510.26787cs.LGcs.AI2025-10被引 21

构建远程工作自动化评估基准,实测AI仅能自动化2.5%任务。

Remote Labor Index: Measuring AI Automation of Remote Work

  • 设计跨行业真实项目评测框架,衡量AI端到端执行能力。
  • 顶尖AI agent仅实现2.5%的自动化率,远低于预期。
  • 为劳动力自动化提供可量化的实证依据,适合政策与企业参考。

尽管AI在知识与推理类研究基准上进展迅速,但其经济价值与自动化潜力仍不明确。为此,我们提出远程工作自动化指数(Remote Labor Index, RLI),一个涵盖多行业的现实世界任务基准,用于评估AI代理在实际场景中的端到端表现。实验显示,当前最先进AI代理在RLI上的自动化率仅为2.5%,接近最低水平。该结果为人工智能自动化讨论提供了实证基础,有助于各方跟踪技术影响并主动应对自动化带来的劳动力变革。

原文摘要 · Abstract (English)

AIs have made rapid progress on research-oriented benchmarks of knowledge and reasoning, but it remains unclear how these gains translate into economic value and automation. To measure this, we introduce the Remote Labor Index (RLI), a broadly multi-sector benchmark comprising real-world, economically valuable projects designed to evaluate end-to-end agent performance in practical settings. AI agents perform near the floor on RLI, with the highest-performing agent achieving an automation rate of 2.5%. These results help ground discussions of AI automation in empirical evidence, setting a common basis for tracking AI impacts and enabling stakeholders to proactively navigate AI-driven labor automation.

AI自动化评估基准劳动力影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。