构建40个传染病多变量预测数据集与基准,推动公共卫生决策智能化
EpiCastBench: Datasets and Benchmarks for Multivariate Epidemic Forecasting

- 整合40个跨疾病、跨区域的多变量疫情数据集,统一评估标准
- 覆盖多种时间粒度与稀疏性特征,揭示全球疫情数据结构规律
- 支持15种模型对比,适合流行病学与机器学习交叉研究者使用
公共卫生领域日益依赖数据驱动决策,使疫情预测成为关键研究方向。近年来,多变量预测模型能更有效捕捉复杂时间依赖关系,优于传统独立建模单序列的方法。然而,由于缺乏高质量、多样化且涵盖不同传染病和地理区域的多变量数据基准,稳健的疫情预测方法发展受限。为此,我们提出EpiCastBench,一个大规模基准框架,包含40个经过筛选的(相关)多变量疫情数据集。这些公开数据集覆盖广泛传染病类型,在时间粒度、序列长度和稀疏性方面表现出多样特征。我们分析了这些数据集的全局特征与结构模式。为确保可复现性和公平比较,建立了统一的评估设置,包括一致的预测时长、标准化预处理流程、多样性能指标及统计显著性检验。基于该框架,我们对15种多变量预测模型进行了全面评估,涵盖统计基线至最先进的深度学习与基础模型。所有数据集与代码已公开于Kaggle(https://www.kaggle.com/datasets/aimltsf/epicastbench)和GitHub(https://github.com/aimltsf/EpiCastBench)。
原文摘要 · Abstract (English)
The increasing adoption of data-driven decision-making in public health has established epidemic forecasting as a critical area of research. Recent advances in multivariate forecasting models better capture complex temporal dependencies than conventional univariate approaches, which model individual series independently. Despite this potential, the development of robust epidemic forecasting methods is constrained by the lack of high-quality benchmarks comprising diverse multivariate datasets across infectious diseases and geographical regions. To address this gap, we present EpiCastBench, a large-scale benchmarking framework featuring 40 curated (correlated) multivariate epidemic datasets. These publicly available datasets span a wide range of infectious diseases and exhibit diverse characteristics in terms of temporal granularity, series length, and sparsity. We analyze these datasets to identify their global features and structural patterns. To ensure reproducibility and fair comparison, we establish standardized evaluation settings, including a unified forecasting horizon, consistent preprocessing pipelines, diverse performance metrics, and statistical significance testing. By leveraging this framework, we conduct a comprehensive evaluation of 15 multivariate forecasting models spanning statistical baselines to state-of-the-art deep learning and foundation models. All datasets and code are publicly available on Kaggle (https://www.kaggle.com/datasets/aimltsf/epicastbench) and GitHub (https://github.com/aimltsf/EpiCastBench).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。