arXiv:2511.19272cs.LG2025-11

2300万参数模型,单卡一周训练,性能超越多数大模型。

Tiny-TSM: Efficiently Training a Lightweight SOTA Time Series Foundation Model

  • 用合成数据生成与增强技术,无需调参即可高效训练。
  • 2300万参数在多任务上达顶尖表现,长短期预测均领先。
  • 新归一化方案加速收敛,适合资源有限场景部署。

我们提出Tiny-TSM,一个参数量仅2300万、训练成本低且性能顶尖的时间序列基础模型。其在单张A100 GPU上不到一周完成训练,依赖全新的合成数据生成与增强流程(SynthTS)。无需神经架构搜索、超参数调优或扩大模型规模,Tiny-TSM在多种时间序列基准数据集上表现卓越,尤其在中长期预测任务中,基于MSE损失的性能全面超越其他已评估的基础模型;短期预测也保持与当前最先进模型相当的水平。此外,我们引入因果输入归一化方案,使模型可采用密集的下一项预测损失进行训练,显著加快收敛速度并缩短训练时长。所有实验均在单张A100 GPU上完成,验证了该方法在资源受限环境中的实用性。

原文摘要 · Abstract (English)

We present Tiny-TSM, a time series foundation model characterized by small scale, economical training, and state-of-the-art performance. It comprises 23M total parameters, trained on a single A100 GPU in less than a week using a new synthetic data generation and data augmentation pipeline (SynthTS). Without any neural architecture search, hyperparameter tuning, or scaling up model size, Tiny-TSM achieves state-of-the-art performance on a wide range of time series benchmark datasets, often outperforming much larger models and even matching the performance of much larger, industrial-scale, likely highly tuned foundation models. Specifically, Tiny-TSM outperforms all other time series foundation models we evaluated on medium- and long-term forecasting tasks under MSE loss, while short-term accuracy is still competitive with state-of-the-art models. We also introduce a causal input normalization scheme that enables time series models to be trained with dense next-token prediction loss, significantly accelerating convergence speed and reducing training time. All experiments were conducted on a single A100 GPU, illustrating the practicality of the proposed approach in a resource-constrained setting.

时间序列轻量模型高效训练基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。