四款时间序列大模型在消费级硬件上实现高精度电力负荷零样本预测。
Time Series Foundation Models for Energy Load Forecasting on Consumer Hardware: A Multi-Dimensional Zero-Shot Benchmark
- 用零样本方式直接预测,无需任务微调,适配低资源环境。
- 长上下文(2048小时)下MASE达0.31,比基准降低47%。
- 适合电网调度、能源管理等需鲁棒性与可解释性的实际场景。
时间序列基础模型(TSFMs)具备零样本预测能力,无需特定任务训练。但其在电力需求预测等关键应用中的准确性、校准性和鲁棒性仍不明确。本研究构建多维度基准,评估四款TSFMs(Chronos-Bolt、Chronos-2、Moirai-2、TinyTimeMixer)及Prophet、SARIMA、Seasonal Naive三类基准模型,基于2020–2024年ERCOT每小时负荷数据,在消费级硬件(AMD Ryzen 7,16GB RAM,无GPU)上运行。评估涵盖四个维度:(1)上下文长度敏感性(24–2048小时),(2)概率预测校准性,(3)分布偏移下的鲁棒性(如新冠疫情封锁与冬季风暴Uri期间),(4)面向运营决策的预测性分析。最优模型在长上下文(2048小时)下日前瞻预测达到近MASE 0.31,较季节性朴素基线降低47%。Prophet在短上下文(24小时)时表现差(MASE > 74),而TSFMs因预训练中学习到的时间模式保持稳定;其中Chronos-2预测区间校准良好(95%经验覆盖率对应90%名义水平),而Moirai-2与Prophet均显过度自信(约70%覆盖率)。研究提供实用模型选择指南,并开源完整基准框架以保障复现性。
原文摘要 · Abstract (English)
Time Series Foundation Models (TSFMs) have introduced zero-shot prediction capabilities that bypass the need for task-specific training. Whether these capabilities translate to mission-critical applications such as electricity demand forecasting--where accuracy, calibration, and robustness directly affect grid operations--remains an open question. We present a multi-dimensional benchmark evaluating four TSFMs (Chronos-Bolt, Chronos-2, Moirai-2, and TinyTimeMixer) alongside Prophet as an industry-standard baseline and two statistical references (SARIMA and Seasonal Naive), using ERCOT hourly load data from 2020 to 2024. All experiments run on consumer-grade hardware (AMD Ryzen 7, 16GB RAM, no GPU). The evaluation spans four axes: (1) context length sensitivity from 24 to 2048 hours, (2) probabilistic forecast calibration, (3) robustness under distribution shifts including COVID-19 lockdowns and Winter Storm Uri, and (4) prescriptive analytics for operational decision support. The top-performing foundation models achieve MASE values near 0.31 at long context lengths (C = 2048h, day-ahead horizon), a 47% reduction over the Seasonal Naive baseline. The inclusion of Prophet exposes a structural advantage of pre-trained models: Prophet fails when the fitting window is shorter than its seasonality period (MASE > 74 at 24-hour context), while TSFMs maintain stable accuracy even with minimal context because they recognise temporal patterns learned during pre-training rather than estimating them from scratch. Calibration varies substantially across models--Chronos-2 produces well-calibrated prediction intervals (95% empirical coverage at 90% nominal level) while both Moirai-2 and Prophet exhibit overconfidence (~70% coverage). We provide practical model selection guidelines and release the complete benchmark framework for reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。