arXiv:2606.23633cs.AIecon.GN2026-06

揭示主流AI工作暴露度评分的局限,呼吁研究与政策协同改进。

AI Exposure Scores: what they measure, what they miss, and what comes next

  • 用静态任务覆盖率衡量AI暴露,忽视动态与地域差异
  • 提出五类新方法应对测量盲区,如动态评估与工人中心指标
  • 强调研究者与政策制定者需共同参与,推动更负责任的未来设计

2023年发布的GPTs are GPTs暴露评分已成为未来工作讨论的核心实证依据,其将暴露定义为大语言模型可辅助的职业任务占比。该研究具有方法论贡献,但随着评分被广泛传播,其原始作者指出的局限性常被忽略。两大差距随之扩大:一是静态评分与政策需求间的结构性错位,时间、地理与概念维度的限制在政策分析中叠加;二是研究与政策间的协调缺失——尽管已有动态评估、集成方法、任务框架扩展、以工人为中心的指标及使用数据等五类应对研究,但政策讨论仍依赖静态评分,未采纳更新方法。本文进一步提出应构建事后评估框架,并开展有意识的政治性工作,重新思考值得追求的未来。研究者须建设数据基础设施、采用参与式方法并面向政策制定者写作;政策方需拓宽证据基础、将工人视为知识伙伴,从预测转向准备。更好测量虽重要,但无法单独弥合这一差距。

原文摘要 · Abstract (English)

A set of exposure scores calculated in 2023 has become a central empirical input to the future of work debate. Produced by Eloundou et al. (2023) and referred to here as the GPTs are GPTs scores, they define exposure as the share of occupational tasks a large language model can assist with. This work is a genuine methodological contribution, but as the scores travel from the time and place they were produced, the limitations the authors named do not always travel with them. Two gaps have widened as a result. The first is structural, between what static exposure scores measure and what policy questions actually require. Taking the diffusion of these scores as a case study, we show how their temporal, geographic, and ontological limitations compound in policy-facing analyses, and we survey five families of research responding to these limits: dynamic and benchmark-based measures, ensemble methods, task-framework extensions, worker-centered metrics, and adoption and usage data. The second gap is the one we argue needs more attention: the coordination between researchers and policymakers. The policy-relevant work which ask who is harmed, who benefits, how, and when, continues to reference the static GPTs are GPTs scores without engagement with the methodological updates that would let these questions be answered more reliably. We then ask what additional steps towards navigating uncertainty remain: ex-post frameworks and the deliberate, political work of reimagining what futures are worthy of building towards are. Closing the research-policy gap is a shared task: policymakers must widen their evidence base, engage workers as epistemic partners, and shift from prediction to preparedness; researchers must build data infrastructure, adopt participatory methods, and write with policymakers in mind. Better measurement matters, but it will not close the second gap alone.

AI评估政策研究未来工作研究协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。