arXiv:2511.02458cs.CLcs.CE2025-11中稿 · ICAIF 25被引 6

用虚拟人物角色提示,GPT-4o可精准预测宏观经济,但角色描述实际无增益。

Prompting for Policy: Forecasting Macroeconomic Scenarios with Synthetic LLM Personas

  • 用经济专家角色提示GPT-4o,模拟央行调查预测
  • 模型与人类专家误差相近,2024-2025年仍保持竞争力
  • 角色描述不影响结果,可省略以降本

我们评估基于人物设定的提示是否提升大语言模型在宏观经济预测中的表现。利用来自PersonaHub语料库的2,368个经济学相关人物角色,我们通过提示GPT-4o,在50个季度周期(2013–2025)内复现欧洲央行专业预测者调查(ECB SPF)。对比了四种目标变量(HICP、核心HICP、GDP增长、失业率)和四个预测时长下的模型与人类专家预测表现,并与100个无角色描述的基线预测进行对照,以分离角色提示的影响。主要发现:第一,GPT-4o与人类专家在准确性上极为接近,差异虽统计显著但实际有限;对2024–2025年数据的样本外评估显示,模型在未见事件上仍保持竞争力,尽管与样本内表现有明显差异。第二,消融实验表明角色描述无显著预测优势,说明可省略该组件以降低计算成本而不损失精度。结果表明,只要提供相关上下文数据,GPT-4o可在样本外宏观事件中实现竞争性预测,同时揭示多样化提示生成的预测结果与人类专家群体相比高度同质。

原文摘要 · Abstract (English)

We evaluate whether persona-based prompting improves Large Language Model (LLM) performance on macroeconomic forecasting tasks. Using 2,368 economics-related personas from the PersonaHub corpus, we prompt GPT-4o to replicate the ECB Survey of Professional Forecasters across 50 quarterly rounds (2013-2025). We compare the persona-prompted forecasts against the human experts panel, across four target variables (HICP, core HICP, GDP growth, unemployment) and four forecast horizons. We also compare the results against 100 baseline forecasts without persona descriptions to isolate its effect. We report two main findings. Firstly, GPT-4o and human forecasters achieve remarkably similar accuracy levels, with differences that are statistically significant yet practically modest. Our out-of-sample evaluation on 2024-2025 data demonstrates that GPT-4o can maintain competitive forecasting performance on unseen events, though with notable differences compared to the in-sample period. Secondly, our ablation experiment reveals no measurable forecasting advantage from persona descriptions, suggesting these prompt components can be omitted to reduce computational costs without sacrificing accuracy. Our results provide evidence that GPT-4o can achieve competitive forecasting accuracy even on out-of-sample macroeconomic events, if provided with relevant context data, while revealing that diverse prompts produce remarkably homogeneous forecasts compared to human panels.

宏观预测LLM提示GPT-4o

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。