arXiv:2601.12343econ.EMcs.AI2026-01被引 2

用等效样本量评估大模型预测人类行为的预训练知识

How Well Do LLMs Predict Human Behavior? A Measure of their Pretrained Knowledge

  • 提出等效样本量衡量大模型预训练知识
  • 实证显示部分经济变量预测能力强,部分弱
  • 适合评估大模型在特定任务中的替代价值

大语言模型(LLMs)被越来越多地用于预测人类行为。本文提出一种衡量预训练模型所携带预测知识的方法:等效样本量,即达到与该模型相同预测精度所需的任务特定数据量。通过比较固定LLM在特定领域的预测误差与其在不断增加的领域数据上训练的灵活机器学习模型的误差,来估算这一指标。此外,我们发展了一种新的交叉验证预测误差渐近理论,支持统计推断。最后,将该方法应用于收入动态面板研究(Panel Study of Income Dynamics)。结果发现,某些经济变量的预测信息在大模型中编码丰富,而另一些则较少,表明大模型作为领域特定数据替代品的价值在不同场景中差异显著。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to predict human behavior. We propose a measure for evaluating how much knowledge a pretrained LLM brings to such a prediction: its equivalent sample size, defined as the amount of task-specific data needed to match the predictive accuracy of the LLM. We estimate this measure by comparing the prediction error of a fixed LLM in a given domain to that of flexible machine learning models trained on increasing samples of domain-specific data. We further provide a statistical inference procedure by developing a new asymptotic theory for cross-validated prediction error. Finally, we apply this method to the Panel Study of Income Dynamics. We find that LLMs encode considerable predictive information for some economic variables but much less for others, suggesting that their value as substitutes for domain-specific data differs markedly across settings.

大模型行为预测等效样本量经济学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。