arXiv:2607.14940cs.LGmath.PR2026-07

在存在干扰和隐藏混杂因素的序列数据中,提出高效因果推断方法。

Causal Inference for Sequential Settings under Interference and Latent Confounding

论文配图:Causal Inference for Sequential Settings under Interference and Latent Confounding
图 1 · 摘自论文原文
  • 用伊辛模型建模单位间依赖,结合外部场与低秩因子结构捕捉处理效应与隐藏混杂。
  • 基于最大伪似然估计,在单一样本下实现参数一致性,可准确估计因果效应。
  • 适用于面板数据、流行病学等存在动态干扰的真实场景,如疫苗率对死亡率影响研究。

我们研究在结果干扰下的序列观测设置中的因果推断问题。具体地,考虑在 T 个时间步上,N 个单位的二值结果呈马尔可夫性;每个时间步的结果依赖关系由伊辛模型刻画,同时受外部场影响,该外部场包含处理效应与隐藏混杂因子。类似面板数据文献,隐藏混杂因子具有低秩因子结构。我们的数据是该高维分布的单一样本。为估计感兴趣的因果量,我们提出一种基于最大伪似然估计(MPLE)的计算高效方法来学习模型参数。在温和假设下,我们建立了参数估计的非渐近一致性,并证明从学习到的模型中采样后能忠实估计因果量。通过合成实验和一个真实案例研究——分析美国各县疫苗接种率对新冠死亡率的因果影响——验证了该方法的有效性。

原文摘要 · Abstract (English)

We study causal inference under outcome interference for sequential, observational settings. Specifically, we consider settings where the binary outcomes over N units are Markovian across T time steps. At each time step, the outcomes of N units have dependencies captured through an Ising model; each outcome is also impacted through an external field capturing the effects of its treatment as well as latent confounders. Similar to panel data literature, these latent confounders are modeled to have a low-rank factor structure. Our data is a single sample from this high-dimensional distribution. To estimate causal quantities of interest, we provide a computationally efficient method based on Maximum Pseudo-Likelihood Estimation (MPLE) for learning the model parameters. Under mild assumptions, we establish non-asymptotic consistency for parameter estimation and show this translates to faithful estimation of causal quantities of interest after sampling from the learned model. We demonstrate the efficacy of the method through synthetic experiments as well as a real-world case-study investigating causal effects of vaccine rates on COVID-19 death rates within US counties nationwide.

因果推断序列数据隐藏混杂伊辛模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。