arXiv:2410.10393cs.LGstat.ML2024-10被引 172

首个面向通用时间序列预测模型的综合评估基准,覆盖超14万时序数据。

GIFT-Eval: A Benchmark For General Time Series Forecasting Model Evaluation

  • 构建跨7领域、10频率的23个数据集,含17700万条数据点
  • 提供约2300亿条非泄露预训练数据,支持模型高效训练
  • 评测17种模型,助力未来时间序列基础模型发展

时间序列基础模型在零样本预测中表现优异,能处理多样任务而无需显式训练。然而,其发展受限于缺乏全面的评估基准。为此,我们提出通用时间序列预测模型评估基准GIFT-Eval,旨在推动跨多种数据集的评估。该基准包含23个数据集,涵盖144,000条时间序列和1.77亿个数据点,覆盖七个领域、十种频率,支持多变量输入及从短期到长期的预测任务。为促进基础模型的有效预训练与评估,我们还提供一个约2300亿条数据点的非泄露预训练数据集。此外,我们对17种基线模型进行了全面分析,包括统计模型、深度学习模型和基础模型,并结合基准特性进行定性讨论,覆盖深度学习与基础模型。我们认为这些洞见,连同新标准的零样本时间序列预测基准,将指导未来时间序列基础模型的发展。代码、数据与排行榜可在https://github.com/SalesforceAIResearch/gift-eval 获取。

原文摘要 · Abstract (English)

Time series foundation models excel in zero-shot forecasting, handling diverse tasks without explicit training. However, the advancement of these models has been hindered by the lack of comprehensive benchmarks. To address this gap, we introduce the General Time Series Forecasting Model Evaluation, GIFT-Eval, a pioneering benchmark aimed at promoting evaluation across diverse datasets. GIFT-Eval encompasses 23 datasets over 144,000 time series and 177 million data points, spanning seven domains, 10 frequencies, multivariate inputs, and prediction lengths ranging from short to long-term forecasts. To facilitate the effective pretraining and evaluation of foundation models, we also provide a non-leaking pretraining dataset containing approximately 230 billion data points. Additionally, we provide a comprehensive analysis of 17 baselines, which includes statistical models, deep learning models, and foundation models. We discuss each model in the context of various benchmark characteristics and offer a qualitative analysis that spans both deep learning and foundation models. We believe the insights from this analysis, along with access to this new standard zero-shot time series forecasting benchmark, will guide future developments in time series foundation models. Code, data, and the leaderboard can be found at https://github.com/SalesforceAIResearch/gift-eval .

时间序列基准测试基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。