构建动态干预下真实疫情时间序列的反事实预测基准
Benchmarking Counterfactual Prediction in Epidemic Time Series with Time-Varying Interventions

- 基于真实数据生成150多个县的疫情反事实轨迹
- 支持静态与动态干预、单政策与多政策场景评估
- 揭示现有方法在真实因果推理中的显著差异
深度学习在时间序列因果推断中取得显著进展,但受限于缺乏具有可观测反事实结果的真实基准。现有数据集或依赖无真实反事实的现实观测,或使用简化模拟,无法捕捉复杂因果动态。为此,我们构建了一个大规模基准,用于动态干预下的疫情时间序列反事实预测。该基准支持静态与时间变化的干预措施,以及单政策和多政策设置,可覆盖广泛的因果推断场景。利用基于真实人口、移动、流行病学及政策数据校准的代理模型,我们在超过150个美国县生成了真实反事实轨迹。通过该基准,我们评估了广泛应用和最先进的因果推断方法,揭示出显著性能差异,并凸显真实时间序列因果推理的挑战。
原文摘要 · Abstract (English)
Deep learning has enabled significant advances in time-series causal inference, yet progress remains constrained by the lack of realistic benchmarks with observable counterfactual outcomes. Existing datasets either rely on real-world observations without ground-truth counterfactuals or on simplified simulations that fail to capture complex causal dynamics. To address this gap, we develop a large-scale benchmark for counterfactual prediction in epidemic time series under dynamic interventions. Unlike existing benchmarks, it supports static and time-varying treatments, as well as both single-policy and multi-policy intervention settings, enabling evaluation of causal inference methods across a broad range of causal inference scenarios. Leveraging a calibrated agent-based model grounded in real-world demographic, mobility, epidemiological, and policy data, we generate realistic counterfactual trajectories across more than 150 U.S. counties. Using this benchmark, we evaluate widely used and state-of-the-art causal inference methods, revealing substantial performance differences and highlighting the challenges of realistic time-series causal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。