arXiv:2602.12147cs.LG2026-02中稿 · ICML被引 11

构建新一代时间序列预测基准,解决数据质量与任务设计缺陷。

It's TIME: Towards the Next Generation of Time Series Forecasting Benchmarks

  • 50个新数据集+98项任务,支持零样本评估
  • 引入真实业务场景任务配置,提升实用性
  • 基于模式层面评估,揭示模型通用能力

时间序列基础模型(TSFMs)正在从特定数据集建模转向可泛化的任务评估。然而,现有基准在四个维度存在共性局限:数据组成受限于重复使用的旧数据源、数据完整性不足缺乏严格质量保障、任务设定脱离真实应用场景、分析视角僵化难以揭示通用规律。为此,我们提出TIME,一个面向任务的新一代基准,包含50个全新数据集和98项预测任务,专为严格的零样本TSFM评估设计,杜绝数据泄露。通过结合大语言模型与人工专家,建立人机协同的基准构建流程,确保高数据完整性,并依据真实业务需求重构任务定义,匹配变量可预测性。此外,提出一种新型模式级评估视角,超越依赖静态元标签的常规数据集级评估。利用结构化时间序列特征刻画内在时序特性,提供跨多样模式的模型能力可泛化洞察。我们评估了12个TSFMs,建立多粒度排行榜,支持深入分析与可视化检查。排行榜见:https://huggingface.co/spaces/Real-TSF/TIME-leaderboard。

原文摘要 · Abstract (English)

Time series foundation models (TSFMs) are revolutionizing the forecasting landscape from specific dataset modeling to generalizable task evaluation. However, we contend that existing benchmarks exhibit common limitations in four dimensions: constrained data composition dominated by reused legacy sources, compromised data integrity lacking rigorous quality assurance, misaligned task formulations detached from real-world contexts, and rigid analysis perspectives that obscure generalizable insights. To bridge these gaps, we introduce TIME, a next-generation task-centric benchmark comprising 50 fresh datasets and 98 forecasting tasks, tailored for strict zero-shot TSFM evaluation free from data leakage. Integrating large language models and human expertise, we establish a human-in-the-loop benchmark construction pipeline to ensure high data integrity and redefine task formulation by aligning forecasting configurations with real-world operational requirements and variate predictability. Furthermore, we propose a novel pattern-level evaluation perspective that moves beyond traditional dataset-level evaluations based on static meta labels. By leveraging structural time series features to characterize intrinsic temporal properties, this approach offers generalizable insights into model capabilities across diverse patterns. We evaluate 12 TSFMs and establish a multi-granular leaderboard to facilitate in-depth analysis and visualized inspection. The leaderboard is available at https://huggingface.co/spaces/Real-TSF/TIME-leaderboard.

时间序列基准测试零样本模式评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。