评估预测模型稳定性,帮供应链减少人工干预。
Measuring Time Series Forecast Stability for Demand Planning
- 用固定输入测同一模型多次输出的波动,量化稳定性。
- 集成模型在保持准确率的同时显著提升预测稳定性。
- 适合关注生产系统可信赖性的供需计划人员。
时间序列预测是供应链需求计划的关键第一步。现有模型实验多聚焦于提升预测准确率,但实际生产中,需求规划者更看重预测结果的一致性与稳定性。若输入未明显变化而预测结果大幅波动,将导致大量人工干预,降低对模型的信任。本文研究模型引发的随机性,即在固定输入下,单个模型产生的预测集方差。方差越低,模型越稳定。我们以M5竞赛和Favorita超市销售数据为基准,对Chronos、DeepAR、PatchTST、Temporal Fusion Transformer、TiDE及AutoGluon最佳集成模型进行案例研究。结果显示,集成模型在不显著降低(甚至提升)准确率的前提下,有效增强了稳定性。尽管结论看似合理,本文核心在于呼吁对部署于生产环境的模型进行更多稳定性研究。
原文摘要 · Abstract (English)
Time series forecasting is a critical first step in generating demand plans for supply chains. Experiments on time series models typically focus on demonstrating improvements in forecast accuracy over existing/baseline solutions, quantified according to some accuracy metric. There is no doubt that forecast accuracy is important; however in production systems, demand planners often value consistency and stability over incremental accuracy improvements. Assuming that the inputs have not changed significantly, forecasts that vary drastically from one planning cycle to the next require high amounts of human intervention, which frustrates demand planners and can even cause them to lose trust in ML forecasting models. We study model-induced stochasticity, which quantifies the variance of a set of forecasts produced by a single model when the set of inputs is fixed. Models with lower variance are more stable. Recently the forecasting community has seen significant advances in forecast accuracy through the development of deep machine learning models for time series forecasting. We perform a case study measuring the stability and accuracy of state-of-the-art forecasting models (Chronos, DeepAR, PatchTST, Temporal Fusion Transformer, TiDE, and the AutoGluon best quality ensemble) on public data sets from the M5 competition and Favorita grocery sales. We show that ensemble models improve stability without significantly deteriorating (or even improving) forecast accuracy. While these results may not be surprising, the main point of this paper is to propose the need for further study of forecast stability for models that are being deployed in production systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。