用计量框架让大模型可靠分析经济文本,避免误判。
Large Language Models: An Applied Econometric Framework
- 区分预测与估计任务,分别设计防泄露和验证机制
- 无验证样本时,不同模型/提示可导致参数估计差异巨大
- 适合想用大模型做实证经济研究的学者
大型语言模型(LLMs)使研究人员能够以空前规模和极低成本分析文本,重新审视旧问题并探索新课题。本文提供了一个计量经济学框架,用于在两类实证应用中实现这一潜力:对于预测问题——从文本中预测结果——必须确保大模型训练数据与研究样本间不存在‘训练泄漏’,这可通过谨慎选择模型和研究设计来实现;对于估计问题——自动化测量经济概念以支持后续分析——有效的下游推断需要将大模型输出与一个小规模验证样本结合,才能获得一致且精确的估计。若缺少验证样本,研究者无法评估大模型输出可能存在的误差,因此看似无害的选择(如使用哪个模型、哪种提示词)可能导致截然不同的参数估计结果。合理使用时,大模型是拓展实证经济学边界的强大工具。
原文摘要 · Abstract (English)
Large language models (LLMs) enable researchers to analyze text at unprecedented scale and minimal cost. Researchers can now revisit old questions and tackle novel ones with rich data. We provide an econometric framework for realizing this potential in two empirical uses. For prediction problems -- forecasting outcomes from text -- valid conclusions require ``no training leakage'' between the LLM's training data and the researcher's sample, which can be enforced through careful model choice and research design. For estimation problems -- automating the measurement of economic concepts for downstream analysis -- valid downstream inference requires combining LLM outputs with a small validation sample to deliver consistent and precise estimates. Absent a validation sample, researchers cannot assess possible errors in LLM outputs, and consequently seemingly innocuous choices (which model, which prompt) can produce dramatically different parameter estimates. When used appropriately, LLMs are powerful tools that can expand the frontier of empirical economics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。