arXiv:2510.00809cs.LG2025-10被引 2

对比大模型与小模型在持续学习中的遗忘问题,发现小模型用技巧可追平大模型表现。

Foundation vs. Specialized Models: Evaluating Catastrophic Forgetting in Continual Time Series Forecasting

  • 比较基础模型(TimesFM-2.0)与专用模型(SamFormer)在持续训练中的表现
  • 微调虽提升新任务准确率,但引发显著遗忘,大模型更抗遗忘
  • 使用遗忘缓解技术(如DER),小模型性能大幅提升,接近大模型

尽管时间序列基础模型(TSFMs)在零样本任务中表现优异,但其在持续微调下的行为尚不明确。本文首次系统研究了在合成与真实世界能源预测基准上,基于时间序列基础模型(TimesFM-2.0、Chronos-2)与专用模型SamFormer在持续学习中的灾难性遗忘现象。结果表明,尽管微调能提升新任务的准确率,却始终引发遗忘;然而,更大模型具有更强的内在鲁棒性。值得注意的是,采用遗忘缓解技术(如DER)后,小模型获得显著收益,最终可与基础模型性能相当。这表明,在非平稳的真实场景下,大型基础模型高昂的计算成本可能无法被其性能优势所抵消,若搭配有效缓解策略,小型模型同样具备竞争力。

原文摘要 · Abstract (English)

While Time Series Foundation Models (TSFMs) excel in zero-shot tasks, their behavior under continual fine tuning is poorly understood. We present the first systematic study of catastrophic forgetting in TSFMs (TimesFM-2.0, Chronos-2) versus a specialized SamFormer model across synthetic and real-world energy forecasting benchmarks. Our results show that while fine-tuning improves new task accuracy, it consistently triggers forgetting, though larger models exhibit greater inherent robustness. Notably, employing forgetting mitigation techniques such as DER, levels the playing field: it provides disproportionate gains to smaller models, allowing them to match TSFM performance by the end of the continual learning sequence. These findings suggest that in realistic, non-stationary scenarios, the high computational cost of large foundation models may not be justified over smaller models equipped with effective mitigation strategies.

时间序列持续学习遗忘抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。