arXiv:2605.15474cs.IR2026-05被引 2

用真实证据评估岗位受AI影响程度,避免仅靠模型猜测。

Jobs' AI Exposure Should Be Measured from Evidence, Not Model Priors

论文配图:Jobs' AI Exposure Should Be Measured from Evidence, Not Model Priors
图 1 · 摘自论文原文
  • 基于检索增强框架,用新闻和论文做证据判断岗位任务是否受AI影响。
  • 72%以上分歧案例中,有证据的判断更受认可,且与现实使用更吻合。
  • 适合政策制定者、企业人力规划者,避免盲目依赖模型预设结论。

本文主张,岗位受AI影响的程度应基于可验证的实证方法衡量,而非仅依赖大模型的先验判断。当前理论性评估多通过零样本提示分类任务级AI暴露,生成无明确证据、缺乏透明推理链、未经外部验证的标签。此类评估结果影响重大,涉及公共与私营资金分配及劳动者对未来前景的认知。为此,我们提出三项标准:可复现性、外部证据支撑、可审查性。本文构建检索增强框架,对O*NET 30.2中的18,796个职业-任务组合,利用开源推理与指令模型,结合检索到的新闻文章与学术论文摘要作为当前AI能力的证据,标注其是否受AI影响。相比零样本基线,有证据条件在自动与人工评估中均于超过72%的分歧案例中更受青睐,且得分更贴近真实世界中的AI应用情况。研究显示,基于证据的测量更能反映当前AI系统实际可行的能力,而非模型单方面宣称。由于AI能力持续演进,相关评估也需动态更新,不能视为不可更改的既定事实。

原文摘要 · Abstract (English)

This position paper argues that job exposure to AI should be measured with grounded, evidence-based methods, not inferred from LLM priors alone. Current theoretical exposure measures use zero-shot prompting to classify task-level AI exposure, generating labels with no explicit evidence, no transparent chain of reasoning, and no external validation. The stakes of these measurements are too high to rely on such methods, as they influence policy making, where public and private funds are directed, and how workers understand their future prospects. We therefore argue that AI capability claims should meet three standards: reproducibility, external grounding, and inspectability. We propose a retrieval-augmented framework that assigns AI exposure labels to all 18,796 occupation--task pairs in O*NET 30.2, using open-weight reasoning and instruct models with retrieved news articles and academic paper abstracts as evidence of current AI capabilities. Relative to a zero-shot baseline, the grounded condition is preferred in over 72\% of disagreement cases under both automatic and human evaluation, and yields scores that align more closely with observed real-world AI usage. Taken together, these findings suggest that evidence-grounded measurement better captures what current AI systems can plausibly do in practice, rather than what a model asserts without external evidence. Because AI capabilities continue to change, the measurements used to inform policy must evolve with them: theoretical AI exposure scores should be periodically reassessed, not inherited as immutable ground truth.

AI影响评估证据驱动政策参考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。