arXiv:2601.16994cs.LG2026-01

将巴西登革热住院数据从月度提升至周级,助力精准预测模型训练。

A Dataset of Dengue Hospitalizations in Brazil (1999 to 2021) with Weekly Disaggregation from Monthly Counts

  • 用三次样条插值法将月度数据拆解为周级,保持每月总量不变。
  • 在圣保罗州高精度数据验证下,三次样条法误差最小,适合作为标准方法。
  • 包含气候、人口、污染等多维变量,适合用于疫情预测与环境健康研究。

本文介绍并公开发布该数据集(v1.0.0),发布于Zenodo,DOI为10.5281/zenodo.18189192。为提升原始月度数据的时间粒度以支持人工智能在流行病预测中的更高效训练,本数据集整合了巴西各市的登革热住院时间序列,并通过带校正步骤的插值协议将其分解至周级(流行病学周)。通过圣保罗州2024年高分辨率参考数据集(同时提供月度和流行病学周计数)对线性插值、抖动法和三次样条三种策略进行评估,结果表明三次样条插值在拟合度上最优,故被采纳用于生成1999至2021年的周级序列。除住院时间序列外,数据集还包含常用于流行病与环境建模的多种解释变量,如人口密度、CH4、CO2、NO2排放量、贫困与城市化指数、最高温度、月均降水量、最低相对湿度及市镇经纬度,均采用相同时间分解方案以确保多变量兼容性。论文详细说明了数据来源、结构、格式、许可协议、局限性及质量指标(MAE、RMSE、R²、KL、JSD、DTW、KS检验),并提供了用于多变量时间序列分析、环境健康研究及机器学习/深度学习爆发预测模型开发的使用建议。

原文摘要 · Abstract (English)

This data paper describes and publicly releases this dataset (v1.0.0), published on Zenodo under DOI 10.5281/zenodo.18189192. Motivated by the need to increase the temporal granularity of originally monthly data to enable more effective training of AI models for epidemiological forecasting, the dataset harmonizes municipal-level dengue hospitalization time series across Brazil and disaggregates them to weekly resolution (epidemiological weeks) through an interpolation protocol with a correction step that preserves monthly totals. The statistical and temporal validity of this disaggregation was assessed using a high-resolution reference dataset from the state of Sao Paulo (2024), which simultaneously provides monthly and epidemiological-week counts, enabling a direct comparison of three strategies: linear interpolation, jittering, and cubic spline. Results indicated that cubic spline interpolation achieved the highest adherence to the reference data, and this strategy was therefore adopted to generate weekly series for the 1999 to 2021 period. In addition to hospitalization time series, the dataset includes a comprehensive set of explanatory variables commonly used in epidemiological and environmental modeling, such as demographic density, CH4, CO2, and NO2 emissions, poverty and urbanization indices, maximum temperature, mean monthly precipitation, minimum relative humidity, and municipal latitude and longitude, following the same temporal disaggregation scheme to ensure multivariate compatibility. The paper documents the datasets provenance, structure, formats, licenses, limitations, and quality metrics (MAE, RMSE, R2, KL, JSD, DTW, and the KS test), and provides usage recommendations for multivariate time-series analysis, environmental health studies, and the development of machine learning and deep learning models for outbreak forecasting.

数据集登革热时间序列预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。