模型输入重构时,不确定性应融入模型而非事后处理。
Learning-based Probabilistic Load Forecasting with Post-hoc and In-model Uncertainty

- 统一框架同步处理时间分辨率与缺失特征重建
- TFT模型结合量化学习使预测区间窄5倍且更准确
- 输入重构误差不可忽视,后验方法会低估风险
智能建筑负荷预测器通常在密集的多变量高频数据上离线训练,但部署时仅能提供小时级、特征受限的输入。缺失特征需重建,其误差会传递至模型。若不反映输入不确定性,预测区间可能失准,影响需求响应调度。本文研究推理输入重建后不确定性应置于何处。提出统一的一日 ahead 概率预测框架,对齐时间分辨率、重建缺失输入并提取因果特征,对比模块化后处理残差分位数方案与集成式模型内分位数学习方案。使用三种中等规模深度学习骨干:循环、混合循环与基于注意力的时序融合变压器(TFT),在相同输入、预测周期、预处理规则和训练预算下进行比较。结果表明,不确定性放置策略依赖于模型结构:TFT采用集成分位数学习最可靠,标签测试窗口上实现2.2-3.6% MAPE与28-83W RMSE,且预测区间宽度约为模块化方案的1/5,在接近名义覆盖率水平下表现更优。Diebold-Mariano检验支持TFT排名及循环模型混合行为。重建敏感性测试显示,重建输入使分位数评分(QS)上升106%,而区间宽度基本不变,说明模型无法自动吸收重建引入的不确定性。对非深度学习基线与季节性留出周的鲁棒性检验也支持该结论。结果揭示了当推理依赖重建输入时,后验残差分位数存在局限。
原文摘要 · Abstract (English)
Smart-building load forecasters are often trained offline on dense, multivariate, high-frequency data, but deployment may provide only hourly, feature-limited inputs. Missing features must then be reconstructed, and their errors can propagate through the model. If this input uncertainty is not reflected, prediction intervals may become miscalibrated, affecting demand-response scheduling. Our work examines where uncertainty should be placed once inference inputs are reconstructed. We develop a unified one-day-ahead probabilistic forecasting framework that aligns temporal resolution, reconstructs the unavailable inputs, and derives causal features, and we compare a modular post-hoc residual-quantile scheme with an integrated in-model quantile-learning scheme. The comparison uses three mid-scale Deep Learning (DL) backbones: recurrent, hybrid recurrent, and attention-based Temporal Fusion Transformer (TFT) models, under identical inputs, forecasting horizon, preprocessing rules, and training budgets. Results show that uncertainty placement is backbone-dependent. Integrated quantile learning is most reliable with the TFT, yielding 2.2-3.6% MAPE and 28-83W RMSE on the labeled test window, while producing intervals about 5x narrower than the modular intervals at the closest-to-nominal coverage level. Diebold-Mariano tests support the TFT ranking and the mixed behavior of the recurrent backbones. A reconstruction-sensitivity test shows that reconstructed inputs increase the Quantile Score (QS) by 106% while interval width remains nearly unchanged, indicating that the model does not automatically absorb reconstruction-induced uncertainty. Robustness checks against non-DL baselines and seasonal hold-out weeks support this ranking. Our results expose the limits of post-hoc residual quantiles when inference depends on reconstructed inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。