利用因果图结构,用改进的期望最大化算法提升域偏移下的目标变量预测准确率。
An Expectation-Maximization Algorithm for Domain Adaptation in Gaussian Causal Models
- 基于因果图构建统一的梯度化EM框架,用投影梯度替代高成本步骤。
- 在协变量偏移和局部机制变化下实现几何收敛,参数误差有理论保证。
- 只重估受域偏移影响的条件分布,适合高维数据,适用于生物网络等复杂场景。
我们研究在部署域发生系统性偏移时,如何对缺失的目标变量进行推断,前提是在源域中存在一个完全可观测的高斯因果有向无环图(Gaussian Causal DAG)。本文提出一种统一的基于期望最大化(EM)的框架,通过因果图结构融合源域与目标域数据,将观测变量的信息迁移至缺失的目标变量。方法上,我们在因果图参数空间中定义了群体级的EM算子,并引入一阶(梯度)更新机制,以单步投影梯度替代耗时的广义最小二乘M步。在标准局部强凹性与光滑性假设,以及类BWY的梯度稳定性(有限缺失信息)条件下,证明该一阶EM算子在真实目标参数附近是局部压缩的,从而在高斯结构方程模型(SEM)下实现几何收敛,并给出参数误差和诱导目标推断误差的有限样本保证,适用于协变量偏移和局部机制偏移。算法上,利用已知因果图冻结源域不变机制,仅重新估计受偏移直接影响的条件分布,使方法可扩展至高维模型。在包含7节点合成模型、64节点MAGIC-IRRI遗传网络及Sachs蛋白信号数据的实验中,所提的因果图感知的一阶EM算法在显著域偏移下,显著优于仅拟合源域的贝叶斯网络和Kiiveri风格的EM基线方法。
原文摘要 · Abstract (English)
We study the problem of imputing a designated target variable that is systematically missing in a shifted deployment domain, when a Gaussian causal DAG is available from a fully observed source domain. We propose a unified EM-based framework that combines source and target data through the DAG structure to transfer information from observed variables to the missing target. On the methodological side, we formulate a population EM operator in the DAG parameter space and introduce a first-order (gradient) EM update that replaces the costly generalized least-squares M-step with a single projected gradient step. Under standard local strong-concavity and smoothness assumptions and a BWY-style \cite{Balakrishnan2017EM} gradient-stability (bounded missing-information) condition, we show that this first-order EM operator is locally contractive around the true target parameters, yielding geometric convergence and finite-sample guarantees on parameter error and the induced target-imputation error in Gaussian SEMs under covariate shift and local mechanism shifts. Algorithmically, we exploit the known causal DAG to freeze source-invariant mechanisms and re-estimate only those conditional distributions directly affected by the shift, making the procedure scalable to higher-dimensional models. In experiments on a synthetic seven-node SEM, the 64-node MAGIC-IRRI genetic network, and the Sachs protein-signaling data, the proposed DAG-aware first-order EM algorithm improves target imputation accuracy over a fit-on-source Bayesian network and a Kiiveri-style EM baseline, with the largest gains under pronounced domain shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。