用大模型解决能源时间序列预测的通用性难题,降低数据与维护成本。
FETS Benchmark: Foundation Models Enable Scalable and Generalizable Energy Time Series Forecasting
- 构建涵盖54个数据集的FETS基准,覆盖9类能源数据与多维使用场景
- 大模型零样本预测表现最优,Chronos-2中位NRMSE达0.472,优于传统模型
- 仅需少量数据、低算力即可部署,适合跨领域快速应用
为应对气候中和能源系统转型需求,精准能源时间序列预测对规划与运行至关重要。然而现有方法依赖特定数据集,训练数据要求高,难以扩展且开发维护成本大。近期预训练大模型在多种预测任务中表现出色,但在能源预测领域尚缺乏系统性评估。本文提出基础模型能源时间序列预测(FETS)基准,包含:(1)从利益相关方、属性和数据类别三维度梳理典型使用场景;(2)基于典型需求收集54个数据集,覆盖9类数据;(3)在不同预测设置下对比基础模型与任务专用机器学习模型。结果表明,引入协变量的零样本基础模型整体表现最佳,Chronos-2中位NRMSE为0.472,紧随其后的是TiRex-2(0.474),均优于XGBoost(0.611)与随机森林(0.696)。分析显示预测性能与谱熵强相关,性能在一定上下文长度后趋于饱和,且随聚合层级提升而改善,如国家负荷、区域供热与电网数据。总体而言,基础模型以最低中位误差、低数据需求与轻量推理,显著降低开发与维护成本,展现出可扩展、通用的能源预测潜力。
原文摘要 · Abstract (English)
Driven by the transition towards a climate-neutral energy system, accurate energy time series forecasting is critical for planning and operations. Yet, it remains a dataset-specific task, requiring comprehensive training data, limiting scalability, and resulting in high model development and maintenance effort. Recently, foundation models aiming to learn generalizable patterns via extensive pretraining have shown strong performance in multiple prediction tasks. Despite their success and strong potential in energy forecasting, a systematic, use-case-differentiated evaluation is still missing. We address this gap by presenting the Foundation Models in Energy Time Series Forecasting (FETS) benchmark. We (1) provide a structured overview of energy forecasting use cases along three main dimensions, i.e., stakeholders, attributes, and data categories, (2) curate 54 datasets across 9 data categories, guided by typical stakeholder interests, and (3) benchmark foundation models against task-specific machine learning across different forecasting settings. In our benchmark study, covariate-informed zero-shot foundation models perform best in aggregate, with Chronos-2 attaining the lowest overall median NRMSE (0.472), closely followed by TiRex-2 (0.474). Both perform better than XGBoost (0.611) and random forest (0.696), although they were trained task-specifically on the full historic target data. Further analysis reveals a strong correlation between predictive performance and spectral entropy. Performance saturates beyond a certain context length and improves with aggregation level, e.g., for national load, district heating, and power grid data. Overall, with the lowest median error, limited data requirements, and low inference and hardware demands, foundation models reduce development and maintenance effort, emerging as scalable and generalizable energy forecasting solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。