用大模型做预测,结合时间数据与外部工具,提升未来事件判断能力。
LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications
- 分三类架构:独立模型、带工具/检索的增强型、与统计模型混合系统
- 实测发现模型对微小输入变化敏感,部分场景下提升不显著
- 适合金融、医疗、能源等领域预测,需关注评估方法可靠性
大语言模型现可支持融合语言推理与时间序列数据、证据检索、外部工具及迭代预测的预报系统。本文研究基于大模型的预测代理,即语言模型参与生成对未观测未来目标的评分预测。将架构分为三类:独立运行的模型处理编码的时间序列或事件上下文;工具与检索增强型代理引入外部证据;混合系统将大模型与统计或基础模型结合。我们回顾训练方法与评估协议,分析正负证据,包括对微小输入扰动的敏感性、移除大模型后准确率未下降的情况,以及可能因数据污染而非时序推理带来的基准提升。应用涵盖金融、天气、健康、能源和运营领域,并总结常用评估基准与数据集。证据表明,度量标准是核心局限。未来工作需在分布偏移下校准,构建抗污染的实时评估,明确报告成本与准确率,以及处理部署预测与实际结果间的反馈关系。
原文摘要 · Abstract (English)
Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external tools, and iterative prediction. We investigate LLM-based forecasting agents, meaning systems in which a language model contributes to a scored prediction about a future or currently unobserved target. We organize architectures into three groups. Standalone LLM workflows operate on encoded time series or event context. Tool- and retrieval-augmented agents incorporate external evidence. Hybrid systems pair LLMs with statistical or foundation models. We then review training methods and evaluation protocols. We examine negative as well as positive evidence, including sensitivity to small input perturbations, ablations in which the LLM component does not improve accuracy, and benchmark gains that may reflect contamination instead of temporal reasoning. We cover applications in finance, weather, health, energy, and operations, and we summarize the benchmarks and datasets used for evaluation. The evidence indicates that measurement is a central limitation. Future work requires calibration under distribution shift, contamination-resistant live evaluation, explicit reporting of cost and accuracy together, and methods for handling feedback between deployed forecasts and the outcomes being forecast.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。