arXiv:2608.30110cs.CLcs.AI2026-09

用大模型实时预测经济指标,效果接近专业机构。

Can LLMs Take the Pulse of the Economy? A Real-Time Evaluation of LLM Nowcasts on Macroeconomic Indicators

论文配图:Can LLMs Take the Pulse of the Economy? A Real-Time Evaluation of LLM Nowcasts on Macroeconomic Indicators
图 1 · 摘自论文原文
  • 让大模型结合网络搜索,每小时更新16个美国经济指标的预估值。
  • 六个月内平均准确率与美联储和彭博专业共识相当。
  • 适合关注实时经济动态的金融从业者和政策研究者。

实时预测宏观经济指标(如GDP、CPI)对货币政策和金融市场至关重要。传统上由央行专家团队进行估算,但效率有限。本文提出LiveMacroEval基准,通过实时生成16个主要美国宏观经济指标的小时级预估值,避免历史数据污染问题。评估采用实时股价波动(LiveMacro Score)和模拟市场交易表现(LiveBetting Score),对比美联储地区银行、彭博ECOS专业共识及auto-ARIMA基线。六个月内,四个具备网络搜索能力的先进大模型在整体准确率上与机构基准相当,但在不同指标间表现差异显著。结果表明,大模型具备作为实时经济状况估计工具的潜力。

原文摘要 · Abstract (English)

Nowcasting headline macroeconomic indicators, i.e., estimating an indicator's value for the current reference period before its official release, is critical for monetary policy and financial markets, and central banks devote dedicated teams of expert economists to producing such estimates. Large language model (LLM) agents are a promising candidate for this task, combining broad world knowledge with real-time web search and supporting queries at higher frequency than institutional nowcasts. Evaluating their nowcasting capability is, however, challenging: headline indicators such as GDP and CPI are widely reported and likely memorized during pretraining, so any evaluation on historical releases is vulnerable to data contamination. To address this, we introduce LiveMacroEval, a live, contamination-resistant benchmark in which LLM agents produce hourly nowcasts for sixteen major U.S. macroeconomic indicators over a pre-release window closing at each official release. Nowcast quality is assessed through a LiveMacro Score against announcement-window equity returns and a LiveBetting Score from simulated Polymarket-style trading, with Federal Reserve regional-bank nowcasts, the Bloomberg ECOS professional consensus, and an auto-ARIMA baseline as comparators. Over six months with four state-of-the-art LLM agents configured with web search, aggregate nowcast accuracy is broadly comparable to the institutional and professional benchmarks, with performance varying widely across individual indicators. This highlights LLM agents' potential as real-time estimators of macroeconomic conditions.

大模型经济预测实时分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。