arXiv:2410.11802cs.LG2024-10KDD被引 68

构建统一基准,评估时间序列大模型的泛化能力

TSFM-Bench: A Comprehensive and Unified Benchmark of Foundation Models for Time Series Forecasting

  • 设计覆盖多场景的评测框架,支持零样本、少样本和全量训练
  • 在跨领域数据集上验证模型表现,揭示现有模型的局限性
  • 提供标准化流程,适合研究者对比新模型与评估方法

时间序列预测(TSF)在金融投资、气象服务和能源管理等领域至关重要。尽管现有方法能力不断提升,但多数需针对特定领域收集数据并重新训练,泛化能力差。基于大规模异构时间序列数据预训练的时间序列基础模型(TSFMs)旨在突破此瓶颈。本研究提出一个综合性统一评测基准TSFM-Bench,用于全面评估各类TSFMs,包括基于大语言模型和专用于时间序列预训练的模型。该基准支持零样本、少样本和全量样本等多种预测场景,涵盖从数据划分、加载、归一化到少样本采样的标准化实验协议,保障评估一致性与公平性。我们在横跨多个领域的多样化数据集上对多种TSFMs进行了广泛评估,揭示了现有模型的优势与内在局限,并为未来模型设计提供了潜在方向。

原文摘要 · Abstract (English)

Time Series Forecasting (TSF) is key functionality in numerous fields, such as financial investment, weather services, and energy management. Although increasingly capable TSF methods occur, many of them require domain-specific data collection and model training and do not generalize well when applied in other domains. Time Series Foundation Models (TSFMs) that are pre-trained on massive heterogeneous time series data aim to overcome these limitations. The prospects for generalizability have spurred the development of a new generation of TSFMs. This study proposes a benchmark, TSFM-Bench, to facilitate comprehensive and unified evaluation of TSFMs. TSFM-Bench covers a wide range of TSFMs, including those based on large language models and those pre-trained on time series data. TSFM-Bench supports multiple forecasting scenarios, including zero-shot, few-shot, and full-shot, enabling assessment across the full range of adaptation strategies. TSFM-Bench also provides a standardized experimental protocols for critical evaluation processes such as dataset splitting, loading, normalization, and few-shot sampling, facilitating consistency and fairness. We report on an extensive evaluation of TSFMs across a diverse range of datasets spanning multiple domains and exhibiting varied statistical characteristics. Specifically, we identify pros and cons and inherent limitations of existing TSFMs, and we propose potential directions for new model designs.

时间序列基础模型评测基准泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。