arXiv:2603.26017cs.LG2026-03被引 5

构建首个覆盖八大时序特性的高质量开源预测基准,推动算法评估标准化。

QuitoBench: A High-Quality Open Time Series Forecasting Benchmark

  • 基于支付宝千亿级流量数据构建跨领域时序基准,按趋势/周期/可预测性划分8类场景。
  • 深度模型在短上下文(L=96)表现更优,而大模型在长上下文(L≥576)胜出,差距达3.64倍MAE。
  • 小模型用59倍少参数达到大模型水平,数据量比模型规模更重要。

时间序列预测在金融、医疗和云计算等领域至关重要,但进展受限于大规模高质量基准的缺乏。为此,我们提出 extsc{QuitoBench},一个涵盖八种趋势×季节性×可预测性(TSF)组合的均衡型基准,聚焦于预测相关特性而非应用领域标签。该基准基于 extsc{Quito}——一个来自支付宝九个业务领域的百亿级时序数据集。在232,200个实例上对10种模型(深度学习、基础模型与统计基线)进行评估,发现:(i) 上下文长度存在交叉点:深度学习模型在短上下文(L=96)占优,而基础模型在长上下文(L≥576)主导;(ii) 可预测性是主要困难来源,导致不同场景间MAE相差3.64倍;(iii) 深度学习模型仅需59倍更少参数即可匹敌或超越基础模型;(iv) 对两类模型而言,增加训练数据量带来的提升远大于扩大模型规模。这些发现经跨基准与跨指标验证具有一致性。本研究开源发布,支持可复现、情境感知的时序预测研究。

原文摘要 · Abstract (English)

Time series forecasting is critical across finance, healthcare, and cloud computing, yet progress is constrained by a fundamental bottleneck: the scarcity of large-scale, high-quality benchmarks. To address this gap, we introduce \textsc{QuitoBench}, a regime-balanced benchmark for time series forecasting with coverage across eight trend$\times$seasonality$\times$forecastability (TSF) regimes, designed to capture forecasting-relevant properties rather than application-defined domain labels. The benchmark is built upon \textsc{Quito}, a billion-scale time series corpus of application traffic from Alipay spanning nine business domains. Benchmarking 10 models from deep learning, foundation models, and statistical baselines across 232,200 evaluation instances, we report four key findings: (i) a context-length crossover where deep learning models lead at short context ($L=96$) but foundation models dominate at long context ($L \ge 576$); (ii) forecastability is the dominant difficulty driver, producing a $3.64 \times$ MAE gap across regimes; (iii) deep learning models match or surpass foundation models at $59 \times$ fewer parameters; and (iv) scaling the amount of training data provides substantially greater benefit than scaling model size for both model families. These findings are validated by strong cross-benchmark and cross-metric consistency. Our open-source release enables reproducible, regime-aware evaluation for time series forecasting research.

时间序列基准测试预测大数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。