用大模型自动识别异常预测,提升零售业预测系统可靠性。
The Forecast Critic: Leveraging Large Language Models for Poor Forecast Identification
- 利用大模型的常识与推理能力,自动检测时间序列预测中的不合理之处。
- 最佳模型F1达0.88,能有效识别趋势错位、峰值异常等错误。
- 无需领域微调即可在真实数据集上发现偏差超10%的不准确预测。
监控预测系统对大型零售企业的客户满意度、盈利能力和运营效率至关重要。我们提出The Forecast Critic,一个利用大语言模型(LLMs)进行自动化预测监控的系统,充分利用其广泛的世界知识和强推理能力。为此,我们系统评估了LLMs评估时间序列预测质量的能力,重点关注三个问题:(1)LLMs能否用于预测监控并识别明显不合理的预测?(2)LLMs能否有效结合非结构化外生特征以判断合理预测应是什么样?(3)性能如何随模型规模和推理能力变化,涵盖当前最先进的LLMs?我们进行了三项实验,包括合成数据和真实世界预测数据。结果表明,LLMs可可靠地检测并批评不良预测,如存在时间错位、趋势不一致和突发误差。我们评估的最佳模型达到F1分数0.88,略低于人类水平(F1: 0.97)。我们还证明,多模态LLMs能有效整合非结构化上下文信号以改进评估;在提供历史促销信息的情况下,模型正确识别缺失或虚假促销峰值的F1为0.84。最后,这些技术在真实世界的M5时间序列数据集上成功识别出不准确预测,不合理预测的sCRPS至少比合理预测高10%。这些发现表明,即使未经领域微调,大模型也可能为自动化预测监控与评估提供可行且可扩展的方案。
原文摘要 · Abstract (English)
Monitoring forecasting systems is critical for customer satisfaction, profitability, and operational efficiency in large-scale retail businesses. We propose The Forecast Critic, a system that leverages Large Language Models (LLMs) for automated forecast monitoring, taking advantage of their broad world knowledge and strong ``reasoning'' capabilities. As a prerequisite for this, we systematically evaluate the ability of LLMs to assess time series forecast quality, focusing on three key questions. (1) Can LLMs be deployed to perform forecast monitoring and identify obviously unreasonable forecasts? (2) Can LLMs effectively incorporate unstructured exogenous features to assess what a reasonable forecast looks like? (3) How does performance vary across model sizes and reasoning capabilities, measured across state-of-the-art LLMs? We present three experiments, including on both synthetic and real-world forecasting data. Our results show that LLMs can reliably detect and critique poor forecasts, such as those plagued by temporal misalignment, trend inconsistencies, and spike errors. The best-performing model we evaluated achieves an F1 score of 0.88, somewhat below human-level performance (F1 score: 0.97). We also demonstrate that multi-modal LLMs can effectively incorporate unstructured contextual signals to refine their assessment of the forecast. Models correctly identify missing or spurious promotional spikes when provided with historical context about past promotions (F1 score: 0.84). Lastly, we demonstrate that these techniques succeed in identifying inaccurate forecasts on the real-world M5 time series dataset, with unreasonable forecasts having an sCRPS at least 10% higher than that of reasonable forecasts. These findings suggest that LLMs, even without domain-specific fine-tuning, may provide a viable and scalable option for automated forecast monitoring and evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。