arXiv:2606.09954cs.LGcs.AI2026-06

不同归一化方法对因果时序大模型的训练与预测影响显著。

Does Normalization Choice Matter for Causal Large Time-Series Models?

论文配图:Does Normalization Choice Matter for Causal Large Time-Series Models?
图 1 · 摘自论文原文
  • 采用基于初始观测值统计的因果归一化,避免未来信息泄露。
  • 实验表明归一化方式直接影响模型收敛速度与预测精度。
  • 适合关注时序建模中训练稳定性与泛化能力的研究者。

面向异构信号集合的大规模时序预测模型正成为主流范式,通常依赖因果自回归架构,逐点基于历史数据进行预测。然而真实世界时序数据普遍存在非平稳性,严重影响预测性能。为缓解此问题,归一化被广泛使用,但在高效因果设置中可能引入未来信息泄漏。尽管已有因果归一化及基于初始观测统计的方法提出,其实际影响仍不明确。本文评估了基于patching和高效因果策略的Transformer型大规模时序模型中的归一化策略,结果表明归一化选择显著影响训练收敛性与预测性能。

原文摘要 · Abstract (English)

Large models for time-series forecasting have been emerged as a promising paradigm for training models on heterogeneous collections of signals. These models typically rely on causal autoregressive architectures, where each observation is sequentially predicted from past. In practice, real-world time-series exhibit non-stationarities, which significantly influence predictive performance. To mitigate this, normalization is commonly employed. However, in efficient causal settings it might induce information leakage from future observations during training. Recent alternatives, including causal normalization and statistics computed from initial observations, have been proposed to address this issue, but their practical implications remain insufficiently understood. In this work, we evaluate normalization strategies for transformer-based large time-series models trained with patching and efficient causal strategy. We showcase that normalization choice significantly influences both training convergence and forecasting performance.

时序模型归一化因果学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。