arXiv:2601.02094cs.LGmath.FA2026-01中稿 · and Presented in I…

提出一种跨模型的时序预测可视化方法,揭示不同模型在时间跨度上的激活模式。

Horizon Activation Mapping for Neural Networks in Time Series Forecasting

  • 基于梯度范数平均构建时间序列激活图,支持多类型神经网络解释
  • 发现批量大小变化呈现指数近似关系,暗示训练动态规律
  • 适用于多变量时序模型对比、验证集选择与模型精细调优

针对时序预测中神经网络模型选择依赖误差指标和特定架构可解释性方法的问题,本文提出一种与模型结构无关的视觉可解释技术——时序激活映射(Horizon Activation Mapping, HAM)。该方法受grad-CAM启发,通过计算每个时间步子序列上梯度更新范数的平均值,分析模型在不同预测时域的响应。引入因果与反因果模式,并定义范数平均均匀分布的等比例线。研究了批量大小、早停、训练/验证/测试划分、模型架构、单变量预测及丢弃率对性能和HAM图的影响。结果显示,不同批量大小下的活动差异呈现出每轮训练间的指数近似趋势。实验使用ETTm2数据集上的MLP-based CycleNet、N-Linear、N-HITS、自注意力FEDformer、Pyraformer、SSM-based SpaceTime以及扩散模型Multi-Resolution DDPM进行绘图分析。其中,N-HITS的神经逼近定理和SpaceTime的指数自回归活动特征均在其训练、验证与测试集的HAM图中得到体现。总体而言,HAM可用于精细化模型选择、验证集设计及跨模型家族比较。

原文摘要 · Abstract (English)

Neural networks for time series forecasting have relied on error metrics and architecture-specific interpretability approaches for model selection that don't apply across models of different families. To interpret forecasting models agnostic to the types of layers across state-of-the-art model families, we introduce Horizon Activation Mapping (HAM), a visual interpretability technique inspired by grad-CAM that uses gradient norm averages to study the horizon's subseries where grad-CAM studies attention maps over image data. We introduce causal and anti-causal modes to calculate gradient update norm averages across subseries at every timestep and lines of proportionality signifying uniform distributions of the norm averages. Optimization landscape studies with respect to changes in batch sizes, early stopping, train-val-test splits, architectural choices, univariate forecasting and dropouts are studied with respect to performances and subseries in HAM. Interestingly, batch size based differences in activities seem to indicate potential for existence of an exponential approximation across them per epoch relative to each other. Multivariate forecasting models including MLP-based CycleNet, N-Linear, N-HITS, self attention-based FEDformer, Pyraformer, SSM-based SpaceTime and diffusion-based Multi-Resolution DDPM over different horizon sizes trained over the ETTm2 dataset are used for HAM plots in this study. NHITS' neural approximation theorem and SpaceTime's exponential autoregressive activities have been attributed to trends in HAM plots over their training, validation and test sets. In general, HAM can be used for granular model selection, validation set choices and comparisons across different neural network model families.

时序预测可解释性神经网络可视化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。