arXiv:2608.18675cs.LG2026-08

对比九种深度模型,发现扩展历史数据能提效但有极限。

An Empirical Benchmark of Deep Time-Series Models for Smart Meter Energy Forecasting

论文配图:An Empirical Benchmark of Deep Time-Series Models for Smart Meter Energy Forecasting
图 1 · 摘自论文原文
  • 用九类模型在两个真实电表数据集上实证对比。
  • 历史窗口越长越准,但超过阈值后提升微弱;预测越远越不准。
  • 轻量模型性能接近复杂模型,适合资源受限场景。

准确预测能源消耗对电力系统高效运行至关重要,直接影响运营成本、能源管理与系统维护。得益于智能电表提供的高分辨率用电数据,数据驱动方法已被广泛用于短期和长期预测。然而,这些方法在真实电表数据上的相对表现仍缺乏充分研究。本文针对九种现代深度学习时序模型(包括线性、MLP、卷积与Transformer架构)构建了实证基准,评估其在两个公开电表数据集上的表现。分析聚焦三个关键因素:历史输入长度、预测时长与模型架构选择。结果表明,延长历史上下文可提升精度,但存在饱和点,超出后收益有限;预测时长越长,精度越低。同时考察了预测精度与计算复杂度的权衡,并评估了模型间差异的统计显著性与实际影响。结果显示,深度学习模型始终优于传统基线,而轻量级架构在显著更低的计算开销下达到相近性能。架构差异仅在长时预测和异质性更高的数据集上产生明显影响。进一步分群分析显示,对多数人口群体而言,模型选择影响有限。

原文摘要 · Abstract (English)

Accurate forecasting of energy consumption is important for the efficient operation of power systems, with direct implications for operational costs, energy management, and system maintenance. Due to the availability of extensive high-resolution consumption data from smart meters, data-driven methods have been used for short-term and long-term forecasting. However, their comparative performance on real-world smart meter data is still not well studied. In this paper, we present an empirical benchmark of nine modern deep learning models for time-series forecasting, including linear, MLP-based, convolutional, and Transformer architectures. We evaluate these models on two publicly available smart meter datasets. Our analysis focuses on three factors that strongly affect forecasting performance: the length of historical input, the prediction horizon, and the choice of model architecture. We show that extending the historical context improves accuracy, but only up to a saturation point, after which additional input provides limited benefit. In contrast, accuracy decreases as the prediction horizon increases. We also investigate the trade-off between prediction accuracy and computational complexity, and assess the statistical significance and practical magnitude of performance differences across models. Our results show that deep learning models consistently outperform classical baselines, while lightweight architectures achieve relatively similar performance at significantly lower computational cost. Additionally, architectural differences only become meaningful at longer forecasting horizons and on more heterogeneous datasets. Finally, a subgroup analysis across geodemographic and household categories shows that model choice has limited impact for most population segments.

时间序列能源预测深度学习实证研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。