构建大规模重尾时间序列数据集,评估模型在极端波动下的预测能力。
TailedTS: Benchmark Dataset for Heavy-Tailed Time Series Prediction and Periodicity Quantification

- 基于维基百科2024年每小时访问数据,构建重尾零膨胀时序数据集。
- 高流量页面周期性弱于低流量页面,传统模型在高流量类别表现显著下降。
- 提供非高斯损失函数评测标准,适合数字平台流量预测研究者使用。
我们提出TailedTS,一个基于2024年维基百科每小时页面访问量的大规模基准数据集,专为测试在重尾、零膨胀和非高斯条件下的时间序列预测模型而设计。数据集包含约246.9亿条记录,覆盖每月约300万唯一页面,以高效Apache Parquet格式存储。维基百科流量呈现明显幂律分布:约5%页面贡献超70%总访问量,形成自然且严苛的极端波动测试环境,现有基准如M4、M5和UCI电力数据集均缺乏此类特征。TailedTS支持多项研究任务:首先,提出基于稀疏自回归与稀疏性及非负性约束的周期性量化框架,发现高频访问页面的周期结构显著弱于低频页面,对大型数字平台的服务器调度与流量预测具有直接意义;其次,提供基于ℓ₁、Huber、分位数和ℓₚ范数等非高斯损失函数的标准化预测基准,表明基于高斯假设的估计器在高流量类别中性能大幅下降,而鲁棒方法在所有流量尺度上均保持稳定优势。TailedTS已公开发布于https://doi.org/10.5281/zenodo.17070469。
原文摘要 · Abstract (English)
We present TailedTS, a large-scale benchmark dataset derived from Wikipedia hourly page view observations throughout 2024, specifically designed to test time series forecasting models under heavy-tailed, zero-inflated, and non-Gaussian conditions. The dataset comprises approximately 24.69 billion data points spanning roughly 3 million unique Wikipedia pages per month, stored in high-efficiency Apache Parquet format. Wikipedia traffic follows a pronounced power-law distribution where roughly 5% of pages account for over 70% of total page views, creating a natural and rigorous testbed for model robustness against extreme volatility that are absent from or underrepresented in existing benchmarks such as M4, M5, and UCI electricity datasets. TailedTS enables several research tasks. First, we introduce a periodicity quantification framework based on sparse autoregression with sparsity and non-negativity constraints, revealing that frequently-viewed pages exhibit significantly weaker periodic structure than their less-viewed counterparts, showing direct implications for server allocation and traffic forecasting on large digital platforms. Second, we provide standardized prediction benchmarks evaluated under a suite of non-Gaussian loss functions, including $\ell_1$-norm, Huber, quantile, and $\ell_p$-norm losses, demonstrating that standard Gaussian-based estimators degrade substantially on high-volume page categories, while robust alternatives provide consistent gains across all traffic scales. TailedTS is publicly available at https://doi.org/10.5281/zenodo.17070469.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。