重新评估跨频段迁移学习,发现统计模型仍优于主流预测大模型。
A More Realistic Evaluation of Cross-Frequency Transfer Learning and Foundation Forecasting Models
- 构建统一神经预测框架,严格避免数据泄露进行跨频段训练评估
- 统计模型在sCRPS上比大模型高8.2%,MASE高出20%以上
- 合成数据预训练可使大模型精度提升7%,但整体仍落后于统计方法
跨频段迁移学习(CFTL)已成为构建大规模时间序列数据集以预训练基础预测模型(FFM)的流行框架。尽管前景可观,现有基准测试方法未能准确评估其性能,主要问题包括:依赖小规模数据集、样本量计算不当、报告次优统计模型,以及未考虑预训练与测试数据集间的重叠风险。为此,我们重构了广泛采用的神经预测网络,适配CFTL设置;仅使用专有和合成数据进行预训练,并严格防止测试泄漏;在15个大型多样化的公开预测竞赛数据集上进行评估。实证分析表明,统计模型的精度常被低估。显著地,在多个数据集上,统计模型及其集成始终优于现有FFM,sCRPS提升超过8.2%,MASE降低超过20%。然而,合成数据预训练可使FFM精度提升7%。
原文摘要 · Abstract (English)
Cross-frequency transfer learning (CFTL) has emerged as a popular framework for curating large-scale time series datasets to pre-train foundation forecasting models (FFMs). Although CFTL has shown promise, current benchmarking practices fall short of accurately assessing its performance. This shortcoming stems from many factors: an over-reliance on small-scale evaluation datasets; inadequate treatment of sample size when computing summary statistics; reporting of suboptimal statistical models; and failing to account for non-negligible risks of overlap between pre-training and test datasets. To address these limitations, we introduce a unified reimplementation of widely-adopted neural forecasting networks, adapting them for the CFTL setup; we pre-train only on proprietary and synthetic data, being careful to prevent test leakage; and we evaluate on 15 large, diverse public forecast competition datasets. Our empirical analysis reveals that statistical models' accuracy is frequently underreported. Notably, we confirm that statistical models and their ensembles consistently outperform existing FFMs by more than 8.2% in sCRPS, and by more than 20% MASE, across datasets. However, we also find that synthetic dataset pre-training does improve the accuracy of a FFM by 7% percent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。