arXiv:2602.09869cs.LG2026-02

在低信噪比时间序列预测中,双路注意力变压器表现优于传统模型。

Statistical benchmarking of transformer models in low signal-to-noise time-series forecasting

  • 采用时序与跨变量交替自注意力机制,提升多变量预测能力。
  • 在信噪比极低(目标与最优预测相关性仅几百分比)时仍保持优异性能。
  • 动态稀疏化注意力矩阵,增强噪声环境下的泛化能力,结果可解释。

我们研究了在仅有几年日度观测数据的低数据场景下,Transformer架构在多变量时间序列预测中的表现。通过生成具有已知时序和横截面依赖结构的合成过程,并设置不同信噪比,我们进行自助法实验,实现对最优真实预测器的样本外相关性直接评估。结果显示,双路注意力变压器(交替使用时序与跨变量自注意力)在广泛设定下均优于标准基线模型——包括Lasso、Boosting方法和全连接多层感知机,尤其在低信噪比条件下表现突出。我们进一步提出一种训练期间应用于注意力矩阵的动态稀疏化策略,发现其在噪声环境中显著有效,此时目标变量与最优预测器的相关性仅为百分之几。对学习到的注意力模式分析揭示了可解释结构,暗示其与经典回归中的稀疏正则化存在联系,为模型在噪声下的良好泛化提供了洞察。

原文摘要 · Abstract (English)

We study the performance of transformer architectures for multivariate time-series forecasting in low-data regimes consisting of only a few years of daily observations. Using synthetically generated processes with known temporal and cross-sectional dependency structures and varying signal-to-noise ratios, we conduct bootstrapped experiments that enable direct evaluation via out-of-sample correlations with the optimal ground-truth predictor. We show that two-way attention transformers, which alternate between temporal and cross-sectional self-attention, can outperform standard baselines-Lasso, boosting methods, and fully connected multilayer perceptrons-across a wide range of settings, including low signal-to-noise regimes. We further introduce a dynamic sparsification procedure for attention matrices applied during training, and demonstrate that it becomes significantly effective in noisy environments, where the correlation between the target variable and the optimal predictor is on the order of a few percent. Analysis of the learned attention patterns reveals interpretable structure and suggests connections to sparsity-inducing regularization in classical regression, providing insight into why these models generalize effectively under noise.

时间序列Transformer低数据信噪比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。