对比局部与全局模型在间歇性时间序列预测中的表现。
Intermittent time series forecasting: local vs global models
- 用神经网络和梯度提升树构建全局模型,对比传统局部模型。
- TiDE模型在4万+真实数据上表现最佳,计算开销更低。
- Tweedie分布对高分位数预测最准,适合库存安全水平设定。
间歇性时间序列(含大量零值)的预测在供应链中至关重要,因库存策略需概率预测来设定安全库存。传统上采用针对每条序列单独训练的局部模型。近年来,基于大规模时间序列训练的全局模型逐渐流行,多采用神经网络或梯度提升树。本文首次系统比较了最先进的概率性局部与全局模型在间歇性时间序列上的表现。全局模型采用三种适配间歇性数据的分布头:负二项分布、漏斗-偏移负二项分布和Tweedie分布。据我们所知,这是首次将后两者与神经网络结合使用。实验基于五个真实世界数据集,涵盖超过4万条时间序列。结果显示,结构简单的神经网络模型TiDE在所有全局模型中精度最高,且始终优于局部模型,同时计算成本更低;而大型全局模型则计算开销大且表现较差。在分布头中,Tweedie对最高分位数的估计最为准确。
原文摘要 · Abstract (English)
Forecasting intermittent time series, which contain zeros, is a crucial challenge in supply chains as inventory policies require probabilistic forecasts to establish safety levels. Intermittent time series are commonly forecast using local models, trained individually on each time series. In the last years global models, trained on a large collection of time series, have become popular for time series forecasting. Global models are often based on neural networks or gradient boosted trees. We carry out the first study comparing state-of-the-art probabilistic local and global models on intermittent time series. For global models we consider three different distribution heads suitable for intermittent time series: negative binomial, hurdle-shifted negative binomial and Tweedie. To the best of our knowledge, this is the first use of the latter two with neural networks. We perform experiments on five datasets comprising overall more than 40'000 real-world time series. Among global models, TiDE, a simple neural network architecture, achieves the best accuracy; it also consistently outperforms local models and has lower computational requirements. Large global models are instead much more computationally demanding and less accurate. Among the distribution heads, the Tweedie provides the best estimates of the highest quantiles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。