用社交媒体文本+大模型,实现更准的通胀预测
LLM-Powered CPI Prediction Inference with Online Text Time Series
- 用大模型从社交文本提取通胀信号,生成每日通胀代理指标
- 融合月度真实数据与日度文本预测,提升预测精度
- 适合经济研究者、政策制定者及量化分析师参考
通货膨胀率(CPI)预测是经济学中重要但具挑战性的任务,现有方法多依赖低频调查数据。随着大语言模型(LLMs)的发展,利用高频在线文本数据提升CPI预测成为新方向。本文提出LLM-CPI,一种结合在线文本时间序列的基于大模型的CPI预测方法。我们从中国主流社交平台收集大量高频文本,使用ChatGPT和训练好的BERT模型为相关帖子构建连续通胀标签。通过LDA和BERT提取文本嵌入,并建立联合时间序列框架:月度模型采用包含观测CPI、文本嵌入与宏观经济变量的ARX结构;日度模型则基于大模型生成的每日CPI代理值与文本嵌入构建VARX结构。我们推导了该方法的渐近性质,并提供两种预测区间构造方式。仿真与真实数据案例验证了LLM-CPI在有限样本下的性能优势与实际应用价值。
原文摘要 · Abstract (English)
Forecasting the Consumer Price Index (CPI) is an important yet challenging task in economics, where most existing approaches rely on low-frequency, survey-based data. With the recent advances of large language models (LLMs), there is growing potential to leverage high-frequency online text data for improved CPI prediction, an area still largely unexplored. This paper proposes LLM-CPI, an LLM-based approach for CPI prediction inference incorporating online text time series. We collect a large set of high-frequency online texts from a popularly used Chinese social network site and employ LLMs such as ChatGPT and the trained BERT models to construct continuous inflation labels for posts that are related to inflation. Online text embeddings are extracted via LDA and BERT. We develop a joint time series framework that combines monthly CPI data with LLM-generated daily CPI surrogates. The monthly model employs an ARX structure combining observed CPI data with text embeddings and macroeconomic variables, while the daily model uses a VARX structure built on LLM-generated CPI surrogates and text embeddings. We establish the asymptotic properties of the method and provide two forms of constructed prediction intervals. The finite-sample performance and practical advantages of LLM-CPI are demonstrated through both simulation and real data examples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。