arXiv:2602.06909cs.LG2026-02被引 4

通用Transformer在时间序列上表现卓越,无需复杂设计即可实现顶尖零样本预测。

Revisiting the Generic Transformer: Deconstructing a Strong Baseline for Time Series Foundation Models

  • 使用标准分块Transformer和简单训练流程,无需特殊架构改进。
  • 在多维度扩展下性能持续提升,零样本预测达到当前最优水平。
  • 开源模型与详尽实验数据,为后续研究提供可复现基准。

时间序列基础模型的兴起推动了该领域快速发展,但各研究采用的训练方式差异较大,难以区分性能提升是源于架构创新还是数据工程。本文深入探究标准分块Transformer的潜力,证明该通用架构仅通过简单训练协议即可实现顶尖的零样本预测性能。我们开展全面消融实验,涵盖模型规模、数据构成与训练技术,以识别高性能的关键因素。结果表明,该通用架构具备优异可扩展性,且在严格控制变量条件下,提供了跨多个维度的模型扩展实证。我们开源了模型及详细发现,旨在为未来研究建立透明、可复现的基准。

原文摘要 · Abstract (English)

The recent surge in Time Series Foundation Models has rapidly advanced the field, yet the heterogeneous training setups across studies make it difficult to attribute improvements to architectural innovations versus data engineering. In this work, we investigate the potential of a standard patch Transformer, demonstrating that this generic architecture achieves state-of-the-art zero-shot forecasting performance using a straightforward training protocol. We conduct a comprehensive ablation study that covers model scaling, data composition, and training techniques to isolate the essential ingredients for high performance. Our findings identify the key drivers of performance, while confirming that the generic architecture itself demonstrates excellent scalability. By strictly controlling these variables, we provide comprehensive empirical results on model scaling across multiple dimensions. We release our open-source model and detailed findings to establish a transparent, reproducible baseline for future research.

时间序列Transformer零样本预测基准模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。