arXiv:2502.14045cs.LGcs.AI2025-02中稿 · TMLR被引 8

复杂模型未必更优,实验设置微调就能让排行榜翻盘。

There are no Champions in Supervised Long-Term Time Series Forecasting

  • 在14个数据集上对比8个顶尖模型,训练超5000个网络验证性能。
  • 更换评估指标或实验细节,冠军模型可能瞬间变垫底。
  • 提醒研究者重视可复现性,别再盲目堆复杂模型。

近期长时序预测领域涌现出众多复杂的监督学习模型,其性能持续超越以往方法。然而,快速进展背后存在基准测试不一致与报告方式混乱的问题,可能影响结果可靠性。本文对主流基准上表现最优的8个模型进行广泛、系统且可复现的评估,覆盖14个数据集,共训练约5000个网络用于超参数搜索。通过全面分析发现,实验设置或评估指标的细微变化即可颠覆对新模型‘领先’的普遍认知。研究强调应从追求更复杂模型转向改进基准测试流程,推动更严谨、标准化的评估,包括可复现的超参数配置和统计检验,以支持更可信的结论,并为未来研究提供建议。

原文摘要 · Abstract (English)

Recent advances in long-term time series forecasting have introduced numerous complex supervised prediction models that consistently outperform previously published architectures. However, this rapid progression raises concerns regarding inconsistent benchmarking and reporting practices, which may undermine the reliability of these comparisons. In this study, we first perform a broad, thorough, and reproducible evaluation of the top-performing supervised models on the most popular benchmark and additional baselines representing the most active architecture families. This extensive evaluation assesses eight models on 14 datasets, encompassing $\sim$5,000 trained networks for the hyperparameter (HP) searches. Then, through a comprehensive analysis, we find that slight changes to experimental setups or current evaluation metrics drastically shift the common belief that newly published results are advancing the state of the art. Our findings emphasize the need to shift focus away from pursuing ever-more complex models, towards enhancing benchmarking practices through rigorous and standardized evaluations that enable more substantiated claims, including reproducible HP setups and statistical testing. We offer recommendations for future research.

时间序列模型评估可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。