arXiv:2604.01363cs.AIecon.GN2026-04

AI自动化更像持续上涨的潮水,而非突然爆发的海啸。

Crashing Waves vs. Rising Tides: Findings on AI Automation from Thousands of Worker Evaluations of Labor Market Tasks

  • 通过6000多个任务和工人评估,发现AI能力是渐进式提升。
  • 2024年模型完成需人1.5小时的任务成功率约60%,2025年超70%。
  • 若趋势延续,2030年前97%文本任务可由大模型以合格质量完成。

我们将AI自动化视为从骤然爆发(如海浪拍岸)到持续提升(如潮水上涨)的连续体。基于来自美国劳工部O*NET分类体系的6000多个文本类、可由大语言模型处理的任务,以及超过6万次由经验丰富的劳动者完成的评估,我们发现缺乏海浪式突变的证据(与现有观点相反)。相反,趋势性上升是主流。当前,大模型在各类任务中的表现已较高且快速进步:2024年第二季度,模型完成人类需1.5小时的任务,成功率约为60%,到2025年第三季度已突破70%。若当前能力增长趋势持续,前沿大模型将在2030年前以88%-97%的成功率完成绝大多数文本类任务,达到最低可接受质量水平。

原文摘要 · Abstract (English)

We characterize AI automation as a continuum between crashing waves, in which capabilities jump abruptly across narrow task sets, and rising tides, in which capabilities improve continuously and broadly. Using evidence from more than 6,000 text-based, LLM-addressable tasks derived from the U.S. Department of Labor's O*NET taxonomy and over 60,000 evaluations by experienced workers, we find little evidence of crashing waves (contrary to existing views). Instead, rising tides are the primary form of AI progress. AI performance is high and improving rapidly across many tasks. In 2024-Q2, models completed text-based tasks that take humans about 1.5 hours to complete with roughly 60% success, rising above 70% by 2025-Q3. If recent trends in AI capability growth persist, frontier LLMs will be able to complete most text-based tasks at minimally sufficient quality with 88%-97% success by 2030.

AI自动化大模型评估劳动力市场趋势预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。