arXiv:2506.21570cs.CLcs.AI2025-06被引 1

预训练语言模型在时间序列预测中持续优于随机初始化模型。

Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting

  • 用预训练语言模型做时间序列预测,比随机初始化更有效。
  • 低数据情况下,预训练模型验证损失持续下降,而随机模型已收敛。
  • 适合关注迁移学习与高效训练的时序研究者。

近期研究证明,在低数据场景下,将预训练语言模型(LMs)适配用于时间序列预测是有效的。本文通过分析上游微调、时间序列分词器和语言模型规模等设计选择对迁移效果的影响,发现这些因素在低数据条件下显著影响验证损失,存在明确更优的选择。与Hernandez等人(2021)不同,我们观察到,语言模型的验证损失在随机初始化模型收敛后仍持续平稳下降,导致跨所有设计选择均存在不可消失的迁移差距。这些发现不仅有助于理解计算高效的时序训练方法,也为研究模型所利用的数据分布的模态无关特性提供了新路径。

原文摘要 · Abstract (English)

Recent works have demonstrated the effectiveness of adapting pre-trained language models (LMs) for forecasting time series in the low-data regime. We build upon these findings by analyzing the effective transfer from language models to time series forecasting under various design choices including upstream post-training, time series tokenizer and language backbone size. In the low-data regime, these design choices have a significant impact on the validation loss, with clear-cut choices that outperform others. Contrary to Hernandez et al. (2021), we observe that the validation loss of the LMs continues to smoothly decrease long after the validation loss of the randomly initialized models has converged, leading to a non-vanishing transfer gap that holds across design choices. These findings not only help shed light on the effective use of compute-efficient training for time series, but also open the way for the study of modality-agnostic properties of data distributions leveraged by these models.

时间序列迁移学习语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。