arXiv:2509.26468cs.LG2025-09被引 40

构建100个真实时间序列预测任务的基准,支持带协变量的任务评估

fev-bench: A Realistic Benchmark for Time Series Forecasting

  • 涵盖7个领域共100个任务,46个含协变量,更贴近真实场景
  • 采用自助法置信区间进行严谨性能统计,避免随机差异误导结论
  • 配套轻量库fev,可无缝集成现有预测流程,提升复现性

基准质量对时间序列预测的有意义评估和持续进展至关重要,尤其在预训练模型兴起的背景下。现有基准普遍存在领域覆盖有限、忽略真实场景(如含协变量的任务)等问题,其聚合方法常缺乏统计严谨性,难以判断性能差异是真实提升还是随机波动。许多基准缺乏一致的评估基础设施,或过于僵化难以融入现有流程。为此,我们提出fev-bench,一个包含100个预测任务的基准,覆盖七个领域,其中46个包含协变量。为支持该基准,我们引入fev——一个轻量级Python库,专注于预测评估的可复现性与现有工作流的集成。借助fev,fev-bench采用基于自助法的置信区间进行合理聚合,从胜率和技能分数两个维度报告性能。我们在fev-bench上评估了预训练模型、统计模型及基线模型,并识别出未来有前景的研究方向。

原文摘要 · Abstract (English)

Benchmark quality is critical for meaningful evaluation and sustained progress in time series forecasting, particularly with the rise of pretrained models. Existing benchmarks often have limited domain coverage or overlook real-world settings such as tasks with covariates. Their aggregation procedures frequently lack statistical rigor, making it unclear whether observed performance differences reflect true improvements or random variation. Many benchmarks lack consistent evaluation infrastructure or are too rigid for integration into existing pipelines. To address these gaps, we propose fev-bench, a benchmark of 100 forecasting tasks across seven domains, including 46 with covariates. Supporting the benchmark, we introduce fev, a lightweight Python library for forecasting evaluation emphasizing reproducibility and integration with existing workflows. Using fev, fev-bench employs principled aggregation with bootstrapped confidence intervals to report performance along two dimensions: win rates and skill scores. We report results on fev-bench for pretrained, statistical, and baseline models and identify promising future research directions.

时间序列基准测试预测评估可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。