生成可调控时间序列因果数据,助力医疗政策等关键领域预测。
DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series
- 构建连续时间干预窗口与反事实采样机制,支持真实场景模拟。
- 提供10万条轨迹、8种结构的精确真值数据集,含干预与反事实配对。
- 适合研究因果建模、强化学习或需可解释决策的高风险领域。
现有时间序列因果推断基准多为观测型、规模小或领域特定,难以满足医疗、政策评估和气候科学中对干预与反事实分析的需求。我们提出DoTime,一个开源、可扩展且理论严谨的多变量时序结构因果模型(TSCM)生成器,随同dotime PyPI包发布四个冻结评估套件。相比已有工作,它新增连续时间干预窗口、带正性保护的反事实采样模式、制度切换的SCM(严格推广中断时间序列)、构造非平稳动态(参数切换),以及在评估窗口内嵌入趋势与结构性断裂的确定性斜坡和正弦干预。此外,验证其作为因果基础模型先验的有效性。所发布套件包含10万条轨迹的训练规模快照,涵盖八种命名识别结构,每种均有精确真值:同一SCM下的成对干预轨迹,及连续时间套件中的共享噪声反事实。配套提供基线实现与评估框架,并提出可检验假设:相同容量下,基于干预数据训练的模型在方向准确率上显著优于仅用观测数据训练的模型。在每种结构、轨迹长度与种子下进行结构匹配的保留测试,所有情况均显示干预先验拟合网络(PFN)误差差距为正。
原文摘要 · Abstract (English)
Most benchmarks for causal inference over time series are observational, small, or domain-specific, leaving interventional and counterfactual estimation under-served exactly where it matters most, such as in healthcare, policy evaluation, and climate science. We introduce \textbf{DoTime}, an open, scalable, and theoretically grounded generator of multivariate temporal structural causal models (TSCMs) with interventions, released as the \code{dotime} PyPI package together with four frozen evaluation suites. Beyond existing work, it adds capabilities absent from prior generators: continuous-time intervention \emph{windows}, counterfactual sampling modes with a positivity guard, regime-switching SCMs as a strict generalization of interrupted time series, non-stationary dynamics by construction with switching SCM parameters, and deterministic ramp and sinusoidal intervention profiles that place trends and structural breaks \emph{inside} the evaluation window. Moreover, it demonstrates the suitability of the generator as a prior for a causal foundation model reference implementation. The released suites span a training-scale snapshot of $100{,}000$ trajectories and eight named identification structures, each with exact ground truth: paired interventional trajectories from the same SCM throughout, and shared-noise counterfactuals in the continuous-time suite. We ship reference baseline implementations with an evaluation harness, and pose a falsifiable claim: interventional training buys a measurable direction-accuracy advantage over an observational model of identical capacity. It is tested across three training seeds per arm. Under structure-matched evaluation on held-out episodes, the interventional prior-fitted network's (PFN) gap is positive in every structure, trajectory length, and seed tested.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。