arXiv:2510.23358cs.CL2025-10被引 2

测试大模型预测AI对就业影响的能力,发现提示设计显著影响预测效果。

How AI Forecasts AI Jobs: Benchmarking LLM Predictions of Labor Market Changes

  • 构建双数据集基准,评估大模型对职业需求变化的预测能力。
  • 结构化提示提升预测稳定性,角色提示更擅长短期趋势判断。
  • 结果受行业和时间跨度影响大,需针对性提示设计。

人工智能正在重塑劳动力市场,但缺乏系统性工具来预测其对就业的影响。本文提出一个基准,评估大语言模型(LLMs)在预测受AI影响职业需求变化方面的表现。该基准结合美国分行业高频职位发布指数与全球AI采用导致的职业变化预测数据,通过明确的时间划分设计预测任务,减少信息泄露风险。采用多种提示策略(任务引导、角色驱动、混合)在不同模型家族中进行评估,考察定量准确性和时间一致性。结果显示,结构化任务提示显著提升预测稳定性,角色提示在短期趋势上表现更优;但性能在不同行业和预测时长间差异明显,凸显领域感知提示的重要性与严格评估协议的必要性。通过开源基准,旨在推动未来关于劳动预测、提示设计及大模型经济推理的研究。本工作为研究大模型作为劳动力市场预测工具的潜力与局限提供了可复现的实验平台。

原文摘要 · Abstract (English)

Artificial intelligence is reshaping labor markets, yet we lack tools to systematically forecast its effects on employment. This paper introduces a benchmark for evaluating how well large language models (LLMs) can anticipate changes in job demand, especially in occupations affected by AI. Existing research has shown that LLMs can extract sentiment, summarize economic reports, and emulate forecaster behavior, but little work has assessed their use for forward-looking labor prediction. Our benchmark combines two complementary datasets: a high-frequency index of sector-level job postings in the United States, and a global dataset of projected occupational changes due to AI adoption. We format these data into forecasting tasks with clear temporal splits, minimizing the risk of information leakage. We then evaluate LLMs using multiple prompting strategies, comparing task-scaffolded, persona-driven, and hybrid approaches across model families. We assess both quantitative accuracy and qualitative consistency over time. Results show that structured task prompts consistently improve forecast stability, while persona prompts offer advantages on short-term trends. However, performance varies significantly across sectors and horizons, highlighting the need for domain-aware prompting and rigorous evaluation protocols. By releasing our benchmark, we aim to support future research on labor forecasting, prompt design, and LLM-based economic reasoning. This work contributes to a growing body of research on how LLMs interact with real-world economic data, and provides a reproducible testbed for studying the limits and opportunities of AI as a forecasting tool in the context of labor markets.

大模型预测就业趋势提示工程劳动市场

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。