arXiv:2604.06906cs.CLcs.AI2026-04

评估大模型对职业能力的自动化影响,发现数学编程最易被替代。

The AI Skills Shift: Mapping Skill Obsolescence, Emergence, and Transition Pathways in the LLM Era

论文配图:The AI Skills Shift: Mapping Skill Obsolescence, Emergence, and Transition Pathways in the LLM Era
图 1 · 摘自论文原文
  • 构建技能自动化可行性指数SAFI,测试4个主流大模型在263项文本任务上的表现。
  • 数学与编程自动化可行性最高(73.2和71.8),倾听与阅读最低(42.2和45.5)。
  • 超七成AI应用是辅助而非替代,适合政策制定者与职场人士参考。

随着大语言模型重塑全球劳动力市场,政策制定者与劳动者亟需实证数据以判断哪些职业能力最易被自动化。本文提出技能自动化可行性指数(SAFI),基于四个前沿大模型(LLaMA 3.3 70B、Mistral Large、Qwen 2.5 72B、Gemini 2.5 Flash)在涵盖美国劳工部O*NET分类中全部35项技能的263项文本任务上的表现(共1,052次模型调用,0%失败率)。结合Anthropic经济指数中的真实AI采用数据(756个职业,17,998项任务),构建AI影响矩阵,将技能分为四类:高替代风险、需再培训、人机协同、低替代风险。关键发现:(1) 数学(SAFI: 73.2)与编程(71.8)自动化可行性最高,主动倾听(42.2)与阅读理解(45.5)最低;(2) 出现“能力需求倒置”现象——在高AI暴露岗位中,需求最高的技能恰恰是大模型表现最差的;(3) 78.7%的观察到的AI交互属于增强而非替代;(4) 四个模型在技能评分上高度一致(最大差距仅3.6分),表明文本类自动化可行性更依赖技能本身而非模型差异。SAFI衡量的是模型对技能的文本表征表现,而非完整职业执行。所有数据、代码与模型响应均已开源。

原文摘要 · Abstract (English)

As Large Language Models reshape the global labor market, policymakers and workers need empirical data on which occupational skills may be most susceptible to automation. We present the Skill Automation Feasibility Index (SAFI), benchmarking four frontier LLMs -- LLaMA 3.3 70B, Mistral Large, Qwen 2.5 72B, and Gemini 2.5 Flash -- across 263 text-based tasks spanning all 35 skills in the U.S. Department of Labor's O*NET taxonomy (1,052 total model calls, 0% failure rate). Cross-referencing with real-world AI adoption data from the Anthropic Economic Index (756 occupations, 17,998 tasks), we propose an AI Impact Matrix -- an interpretive framework that positions skills along four quadrants: High Displacement Risk, Upskilling Required, AI-Augmented, and Lower Displacement Risk. Key findings: (1) Mathematics (SAFI: 73.2) and Programming (71.8) receive the highest automation feasibility scores; Active Listening (42.2) and Reading Comprehension (45.5) receive the lowest; (2) a "capability-demand inversion" where skills most demanded in AI-exposed jobs are those LLMs perform least well at in our benchmark; (3) 78.7% of observed AI interactions are augmentation, not automation; (4) all four models converge to similar skill profiles (3.6-point spread), suggesting that text-based automation feasibility may be more skill-dependent than model-dependent. SAFI measures LLM performance on text-based representations of skills, not full occupational execution. All data, code, and model responses are open-sourced.

AI影响技能评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。