arXiv:2604.02403econ.EMcs.CL2026-04被引 1

用大模型量化工作认知内容,填补传统调查无法触及的细粒度空白。

Measuring What Cannot Be Surveyed: LLMs as Instruments for Latent Cognitive Variables in Labor Economics

  • 基于语义外生性等四条件,将大模型输出作为认知变量测量工具
  • 构建18,796条任务的劳动技能指数,与多个AI暴露指数高度相关(最高r=0.85)
  • 方法可推广至需大规模语义量化研究的领域,如教育、政策评估

本文建立了使用大语言模型(LLMs)作为潜在经济变量测量工具的理论与实践基础,聚焦于现有调查工具难以实现细粒度刻画的职业任务认知内容。作者形式化了四个有效工具变量条件:语义外生性、构念相关性、单调性与模型不变性。应用该框架构建了基于18,796条O*NET任务描述的增强型人力资本指数(AHC_o),由Claude Haiku 4.5评分,并与六个现有AI暴露指数进行验证。结果表明,该指数具有强收敛效度(与Eloundou GPT-gamma相关系数r=0.85,与Felten AIOE为r=0.79)和区分效度。主成分分析确认AI相关职业指标涵盖两个独立维度——增强与替代。两名大模型在3,666组配对评分中显示互评一致性(皮尔逊相关r=0.76,Krippendorff's alpha=0.71)。四种提示框架的敏感性分析表明任务排序稳健。显然相关的工具变量(ORIV)估计得到的系数比普通最小二乘法(OLS)高25%,符合经典测量误差衰减预期。该方法可推广至任何需要大规模语义内容量化的领域。

原文摘要 · Abstract (English)

This paper establishes the theoretical and practical foundations for using Large Language Models (LLMs) as measurement instruments for latent economic variables -- specifically variables that describe the cognitive content of occupational tasks at a level of granularity not achievable with existing survey instruments. I formalize four conditions under which LLM-generated scores constitute valid instruments: semantic exogeneity, construct relevance, monotonicity, and model invariance. I then apply this framework to the Augmented Human Capital Index (AHC_o), constructed from 18,796 O*NET task statements scored by Claude Haiku 4.5, and validated against six existing AI exposure indices. The index shows strong convergent validity (r = 0.85 with Eloundou GPT-gamma, r = 0.79 with Felten AIOE) and discriminant validity. Principal component analysis confirms that AI-related occupational measures span two distinct dimensions -- augmentation and substitution. Inter-rater reliability across two LLM models (n = 3,666 paired scores) yields Pearson r = 0.76 and Krippendorff's alpha = 0.71. Prompt sensitivity analysis across four alternative framings shows that task-level rankings are robust. Obviously Related Instrumental Variables (ORIV) estimation recovers coefficients 25% larger than OLS, consistent with classical measurement error attenuation. The methodology generalizes beyond labor economics to any domain where semantic content must be quantified at scale.

大模型应用认知测量劳动经济学工具变量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。