arXiv:2605.19231cs.LGstat.ML2026-05

提出DeRegiME模型,用可解释的波动模式分解提升时间序列概率预测精度。

DeRegiME: Deep Regime Mixtures for Probabilistic Forecasting under Distribution Shift

论文配图:DeRegiME: Deep Regime Mixtures for Probabilistic Forecasting under Distribution Shift
图 1 · 摘自论文原文
  • 基于稀疏变分GP与共享门控机制,分离信号与不确定性波动模式。
  • 在10个基准上降低20.3%的NLPD,CRPS和MSE也显著改善。
  • 能识别突发、渐进、季节性等不同分布漂移,适合需要可解释性的场景。

我们提出DeRegiME——一种直接多步概率预测模型,将潜在不确定性波动模式与基础信号分离,并通过稀疏变分高斯过程(GP)中的共享门控机制,将每个预测点软分配到学习到的周期性波动模式中。其非平稳的波动混合核与Student-t似然结合各模式子核与噪声过程,生成单一稀疏GP后验,而非多个专家混合。该模型克服了神经网络预测的局限:点预测忽略残差不确定性,而现有概率头(如单边际、未解释混合、分位数集或扩散样本)通常无法揭示残差中的波动结构。在噪声异方差时间序列中,分布漂移可能突变、渐进或依赖时序跨度,常体现在残差中而非条件均值。DeRegiME提供可解释的均值-残差-噪声分解,以特征空间的直和表示锚定波动模式为残差相似性聚类,其转换表现为隐式变点。有效波动模式数量由截断棒破门控自动剪枝。我们证明了核函数有效性与预测密度合理性。在十个基准与三个编码器网格上,相比最强基线(DeepAR/GluonTS风格动态Student-t头),DeRegiME在负对数预测密度(NLPD)上平均降低20.3%,同时在CRPS(3.0%)和MSE(4.7%)上也有并行提升。改进在所有数据集上保持一致,涵盖突变、渐进与季节性漂移。

原文摘要 · Abstract (English)

We introduce DeRegiME -- Deep Regime Mixture of Experts -- a direct multi-horizon probabilistic forecaster that separates latent uncertainty regimes from the underlying signal and softly assigns each forecast location to learned recurring regimes using a sparse variational Gaussian process (GP) whose nonstationary regime-mixing kernel and Student-t likelihood combine per-regime sub-kernels and noise processes via a shared gate. This yields a single sparse-GP posterior, not a mixture of GP experts. DeRegiME addresses a key limitation of neural forecasters: point forecasts discard residual uncertainty, and probabilistic heads -- whether single marginals, uninterpreted mixtures, quantile sets, or diffusion samples -- rarely expose the regime structure of the residual. Yet distribution shift in noisy heteroskedastic time series may be abrupt, gradual, or horizon-dependent and often appears in residual uncertainty rather than the conditional mean. DeRegiME yields an interpretable mean-residual-noise decomposition with a direct-sum feature-space representation that anchors regimes as clusters of residual similarity whose transitions surface as implicit changepoints. The effective number of regimes is pruned by the stick-breaking gate. We prove kernel validity and predictive-density propriety, and across ten benchmarks and three encoder grids DeRegiME improves negative log predictive density (NLPD) by 20.3% over the strongest encoder-matched baseline, a DeepAR/GluonTS-style dynamic Student-t head, with parallel gains on CRPS (3.0%) and MSE (4.7%). Improvements are consistent across all datasets, which span abrupt, gradual, and seasonal shifts.

概率预测时间序列波动模式可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。