arXiv:2605.00015eess.SPcs.AI2026-05被引 2

用强化学习提升时间序列模型泛化能力,应对数据分布变化挑战。

TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning

论文配图:TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning
图 1 · 摘自论文原文
  • 通过强化学习设计奖励机制,精细化评估每步预测贡献。
  • 在多种数据环境下均优于传统监督微调方法,提升准确率与泛化性。
  • 适合数据稀缺或分布变化大的实际时间序列预测任务。

时间序列基础模型(TSFMs)通过大规模预训练展现出强大的泛化能力和数据效率。然而,在下游任务中适应时仍面临时间分布偏移和数据可用性差异的挑战。由于时间序列具有非平稳性和不确定性,历史训练与未来预测分布存在差异,导致现有基于监督微调(SFT)的方法易过拟合、泛化能力受限。此外,不同任务的数据条件各异,要求模型从有限样本中提取通用的时间模式。为此,我们提出时间序列强化微调(TimeRFT),一种基于强化学习的TSFM适配范式。TimeRFT引入两个面向预测的训练策略:(i) 质量感知的时间奖励机制,通过综合评估每一步预测对整体性能的贡献,实现细粒度的信用分配;(ii) 困难度感知的数据选择策略,优先选取具有可泛化预测模式的信息丰富样本。在多个真实世界预测基准上的实验表明,TimeRFT在不同数据条件下均显著超越SFT方法,提升预测精度并增强对未预见分布偏移的鲁棒性。

原文摘要 · Abstract (English)

Time Series Foundation Models (TSFMs) have demonstrated strong generalization capability and data efficiency in time series forecasting through large-scale pretraining. However, adapting TSFMs to downstream forecasting tasks remains challenging due to temporal distribution shifts and varying data availability. Specifically, the non-stationary and uncertain nature of time series data leads to discrepancies between historical training and future forecasting distributions, making existing Supervised FineTuning (SFT)-based adaptation vulnerable to overfitting and limited generalization. Moreover, forecasting tasks often operate under varying data regimes, requiring TSFMs to extract generalizable temporal patterns from limited training samples. To address these challenges, we propose Time series Reinforcement FineTuning (TimeRFT), a reinforcement learning-based adaptation paradigm for TSFMs. TimeRFT introduces two forecasting-oriented training recipes: (i) A quality-aware temporal reward mechanism providing fine-grained credit assignment by holistically evaluating the contribution of each prediction step to overall forecasting performance. (ii) A difficulty-aware data selection strategy prioritizing informative time series samples with generalizable forecasting patterns. Extensive experiments on diverse real-world forecasting benchmarks demonstrate that TimeRFT consistently surpasses SFT-based adaptation methods across various real-world forecasting tasks with different data regimes, achieving improved prediction accuracy and enhanced generalization against unforeseen distribution shifts.

时间序列强化学习微调泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。